Curie Brief
Turn on cookies to sign in
Signing in saves your progress to your Curie account. We can only do that with cookies on — turn them on to continue.

When it comes to complex coronary artery disease, AI tools like ChatGPT-4o and Gemini 2.0 only align with multidisciplinary heart teams about 75% of the time — and that's only when given rich clinical detail. Feed them less information, and agreement drops to just 35%. Researchers say the findings highlight both AI's promise and its pitfalls as a clinical decision-support tool.
A new single-center study out of the University of Melbourne found that large language models (LLMs) like ChatGPT-4o and Gemini 2.0 show only "moderate" agreement with multidisciplinary heart teams when recommending treatment for complex coronary artery disease (CAD). When given detailed, free-text clinical narratives, the AI tools agreed with the heart team in 75% of cases — reasonable, but far from perfect. Drop the detail level to a structured 30-variable form, and agreement plummeted to just 35%.
The study, published in JSCAI, reviewed 546 patients with complex CAD who had treatment recommendations made by a heart team between 2019 and 2024. Notably, adding guideline-based information to the AI prompts didn't improve performance. Cases where ChatGPT-4o and the heart team disagreed were associated with a higher risk of major adverse cardiovascular events (MACE), though this became non-significant after adjusting for clinical complexity.
Key Takeaways
Why it matters: As patients and clinicians increasingly turn to AI for medical guidance, this study is a timely reminder that AI recommendations can be misleading when based on incomplete information — especially in high-stakes cardiac decisions.