Clinical Expertise Safeguards Against Persuasive AI Errors in Radiology
A South Korean study warns that high-quality AI rationales can deceive clinicians, reinforcing the critical role of human expertise.
Clinical expertise and reader confidence are the most effective defenses against incorrect AI-generated diagnostic suggestions, according to a new retrospective study. The research, published in the RSNA's Radiology journal, warns that persuasive but unsound reasoning from large language models (LLMs) can mislead radiologists, regardless of the model's actual accuracy.
South Korean researchers tasked 10 readers with interpreting chest imaging from 100 patients. The study utilized a dataset from the Korean Society of Thoracic Radiology Weekly Case platform (2018-2020), which included PET scans, MRIs, CTs, and X-rays. Readers compared their own interpretations against suggestions from two different models: a high-accuracy model based on GPT-5 (76% accuracy) and a low-accuracy model based on GPT-4o (27% accuracy).
The findings showed that while model confidence (OR 3.82) and reader expertise (OR 2.06) were independently associated with adequate interaction, the AI's ability to explain its reasoning created a significant risk. Specifically, higher rationale quality increased the acceptance of incorrect suggestions (OR 1.71), indicating that a well-constructed argument can trick a clinician into accepting a wrong answer.
The Paradox of Explainability
This phenomenon highlights a double-edged sword in AI explainability. While the goal of providing rationales is to help clinicians understand the AI's logic, the study suggests that "persuasive" rationales can act as a deceptive tool. When an LLM presents an incorrect answer with high confidence and a polished explanation, it can override a clinician's intuition.
However, the data confirms that human factors provide a necessary check. Reader expertise (OR 0.54) and reader confidence (OR 0.80) acted as protective factors, significantly reducing the likelihood that a radiologist would be fooled by an incorrect AI suggestion.
Implications for Clinical Practice
These results suggest that LLMs do not replace the need for deep clinical knowledge but instead amplify its importance. As AI becomes more integrated into diagnostic workflows, the risk of "automation bias"—where users trust the machine over their own judgment—increases.
The study underscores that the radiologist's role is evolving from a primary interpreter to a critical supervisor who must shape both the inputs and the final interpretation of AI outputs. Without this expert layer, the tendency of LLMs to produce confident but incorrect "hallucinations" could lead to diagnostic errors.
Future Outlook
Expertise serves as the essential safeguard against persuasive but incorrect rationales. The research indicates that the integration of AI in radiology must be accompanied by training that emphasizes the skepticism of AI-generated rationales. Future studies may need to examine whether this susceptibility varies across different medical specialties or if specific types of AI-generated errors are more persuasive than others. For now, the evidence suggests that the most reliable diagnostic loop remains one where the human expert holds the final veto.