{"id":19047,"date":"2026-08-13T05:21:31","date_gmt":"2026-08-13T08:21:31","guid":{"rendered":"https:\/\/rtmedical.com.br\/tmp-en-1786609290674\/"},"modified":"2026-08-13T05:21:38","modified_gmt":"2026-08-13T08:21:38","slug":"radiologists-llm-incorrect-advice","status":"publish","type":"post","link":"https:\/\/rtmedical.com.br\/en\/radiologists-llm-incorrect-advice\/","title":{"rendered":"What Keeps Radiologists From Trusting Bad AI Advice"},"content":{"rendered":"<h2>Expertise is what separates help from trap<\/h2>\n<p>A South Korean study published in <em>Radiology<\/em>, the RSNA&#8217;s flagship journal, identified what determines whether a radiologist gains or loses by consulting a large language model: the confidence the model itself expresses and, above all, the reader&#8217;s expertise in the modality at hand. When the algorithm&#8217;s explanation is well written but the diagnosis is wrong, subspecialized readers resist; those without that background go along. The quality of the reasoning, which ought to be a benefit, cuts both ways.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" class=\"alignleft lazyload\" data-src=\"https:\/\/rtmedical.com.br\/wp-content\/uploads\/2026\/08\/radiologista-llm-sala-de-laudos.jpg\" alt=\"Radiologist interpreting an imaging study on a diagnostic monitor in a dimly lit reading room\" width=\"620\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1880px; --smush-placeholder-aspect-ratio: 1880\/1253;\"><figcaption>Consulting an LLM mid-read changes the decision dynamic \u2014 and not always for the better<\/figcaption><\/figure>\n<h2>How the study was designed<\/h2>\n<p>The retrospective work was led by Taehee Lee and colleagues in the Department of Radiology at Seoul National University Hospital and College of Medicine, published Aug. 11. Ten readers interpreted chest imaging from 100 patients. Cases came from the Korean Society of Thoracic Radiology&#8217;s Weekly Case platform, spanning 2018 to 2020 and including radiographs, CT, MRI and PET.<\/p>\n<p>The design has one elegant detail: each reader went through two sessions. In the first, they interpreted without any assistance. In the second, they received support from one of two randomized models \u2014 a high-accuracy model at 76% correct, built on GPT-5, and a low-accuracy one at 27%, using GPT-4o. Readers were given multiple-choice diagnostic options along with the model&#8217;s rationales, and the reference standard was set by each case&#8217;s author.<\/p>\n<p>Deliberately including a poor model is the smartest part of the protocol. Knowing whether AI helps when it is right is not enough \u2014 the clinically relevant question is what happens when it is wrong with poise. A model with 27% accuracy is wrong in three of every four cases, and still writes with the same fluency as the model that is right in three of four.<\/p>\n<h2>The numbers: who resists and who yields<\/h2>\n<p>Two variables were independently associated with adequate radiologist-model interaction: model confidence (odds ratio 3.82) and reader expertise (OR 2.06). For readers unfamiliar with the term, an odds ratio above 1 indicates increased odds of the outcome; below 1, reduced odds. An OR of 3.82 is a strong effect.<\/p>\n<figure class=\"wp-block-image size-large\"><img decoding=\"async\" class=\"alignright lazyload\" data-src=\"https:\/\/rtmedical.com.br\/wp-content\/uploads\/2026\/08\/radiografia-torax-analise-medica.jpg\" alt=\"Physician examining a chest radiograph with pulmonary findings against the light\" width=\"620\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1880px; --smush-placeholder-aspect-ratio: 1880\/1253;\"><figcaption>One hundred chest imaging cases \u2014 radiography, CT, MRI and PET \u2014 formed the evaluation set<\/figcaption><\/figure>\n<p>The effect of model confidence, however, was weaker among expert readers (OR 0.79). That tracks: the specialist is not impressed by emphasis. They already hold their own hypothesis and use the model&#8217;s output as a counterpoint rather than an anchor.<\/p>\n<p>The most uncomfortable finding concerns rationale quality. Better-constructed arguments reduced rejection of correct suggestions (OR 0.79) \u2014 good \u2014 but increased acceptance of incorrect ones (OR 1.71) \u2014 bad. It is the same mechanism running in opposite directions: a persuasive explanation persuades, regardless of whether it is right. Two factors proved protective against that: reader expertise (OR 0.54) and the reader&#8217;s own confidence in their hypothesis (OR 0.80), both reducing adherence to error.<\/p>\n<h2>What the authors conclude<\/h2>\n<p>The researchers&#8217; phrasing is precise: &#8220;Although high model confidence was associated with correct decisions, expertise may serve as a safeguard against persuasive but incorrect rationales. The double-edged role of rationale quality suggests explainability benefits those able to critically appraise it. Overall, effective LLM assistance depends on factors beyond model performance, including reader expertise and model confidence.&#8221;<\/p>\n<p>They go further, in a passage that runs against the substitution narrative: &#8220;Beyond serving as a safeguard, reader expertise anchors human-LLM interaction. Given current limitations of vision-language models, expert-authored descriptions help guide model reasoning and reduce hallucinations. Expertise also enables clinicians to pose focused questions, evaluate outputs, and distinguish persuasive but unsound reasoning from clinically valid explanations. Rather than diminishing radiologists&#8217; roles, LLMs amplify the importance of clinical expertise, as radiologists shape both inputs and interpretation.&#8221;<\/p>\n<h2>Technical context: why fluency deceives<\/h2>\n<p>The mechanism is worth understanding, because it is not a passing defect of one model version. A language model is optimized to produce plausible text given context \u2014 not to signal its own uncertainty in a calibrated way. Hence the patterns radiologists already recognize in practice: factually incorrect explanations delivered with undue conviction, omission of a relevant finding, pattern overgeneralization from an atypical case.<\/p>\n<p>The trouble is that human perception of competence leans heavily on fluency. Organized prose with correct terminology and visible reasoning structure triggers the same heuristic we use to judge a colleague at a case conference. Except the colleague has a world model behind the words, and the LLM has a probability distribution over tokens. For someone who commands the subject, the difference surfaces at the first anatomical inconsistency; for someone who does not, it never surfaces.<\/p>\n<h2>Practical implications for imaging services<\/h2>\n<p>The operational translation is concrete. First: LLM assistance should not be offered indiscriminately to trainees working outside their area of command. That is precisely the combination with the greatest documented harm \u2014 low expertise plus persuasive rationale.<\/p>\n<p>Second: order matters. Recording your own hypothesis before consulting the model preserves reader confidence, which the study showed to be protective. Consulting first and forming an opinion afterward hands the anchor to the algorithm.<\/p>\n<p>Third: explainability is not a universal safeguard. Industry sells &#8220;explainable AI&#8221; as the answer to the trust problem, and this study shows explanation helps those who can critique it and harms those who cannot. In residency programs, that argues for explicitly training critical appraisal of model outputs, much as we train critical appraisal of journal articles.<\/p>\n<p>There is a direct parallel with the debate over patient communication. The ACR published <a href=\"https:\/\/rtmedical.com.br\/en\/acr-patient-guide-ai-summaries\/\">guidance on AI-generated report summaries<\/a> precisely because fluent text and correct text are not the same thing. And the issue gains urgency knowing that <a href=\"https:\/\/rtmedical.com.br\/en\/patients-read-reports-first\/\">44% of reports reach the patient before the physician<\/a> \u2014 many will be interpreted with a chatbot&#8217;s help, no radiologist in the loop.<\/p>\n<h2>Limits and what to watch next<\/h2>\n<p>This is a retrospective study with ten readers and curated cases from an educational platform \u2014 material selected for teaching value, typically harder and more atypical than a routine worklist. Generalizing to a general service&#8217;s daily volume calls for caution. Patient outcomes were not assessed, only agreement with a reference standard, and the model versions used will age quickly.<\/p>\n<p>Even so, the identified mechanism should survive version churn, because it follows from how these systems are built and how humans judge competence. The practical conclusion for anyone working with <a href=\"https:\/\/rtmedical.com.br\/en\/raidium-ai-oncology-imaging-platform\/\">AI oncology imaging platforms<\/a> or any generative tool in the diagnostic pipeline is less glamorous than the market promise: a tool&#8217;s value is proportional to the competence of whoever operates it. The better the AI argues, the more expert its listener needs to be.<\/p>\n<p><strong>Source:<\/strong> <a href=\"https:\/\/radiologybusiness.com\/topics\/artificial-intelligence\/factors-make-radiologists-less-likely-be-fooled-large-language-models\" target=\"_blank\" rel=\"noopener\">Radiology Business<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A Radiology study shows persuasive LLM rationales raise acceptance of wrong answers 71%, and reader expertise is the antidote. See the odds ratios.<\/p>\n","protected":false},"author":1,"featured_media":19018,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"om_disable_all_campaigns":false,"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"ngg_post_thumbnail":0,"_rt_cluster":"","fifu_image_url":"","fifu_image_alt":"","footnotes":""},"categories":[102,100],"tags":[],"class_list":["post-19047","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","category-radiology"],"aioseo_notices":[],"rt_seo":{"title":"","description":"A Radiology study shows persuasive LLM rationales raise acceptance of wrong diagnoses, while reader expertise protects. See the odds ratios.","canonical":"","og_image":"","robots":"index,follow","schema_type":"Article","include_in_llms":true,"llms_label":"Expertise shields radiologists from LLM errors","llms_summary":"A Seoul National University Hospital study in Radiology with 10 readers and 100 chest imaging cases found higher-quality rationales increase acceptance of incorrect LLM suggestions (OR 1.71), while reader expertise (OR 0.54) and reader confidence (OR 0.80) are protective.","faq_items":[],"video":[],"gtin":"","mpn":"","brand":"","aggregate_rating":[]},"_links":{"self":[{"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/posts\/19047\/"}],"collection":[{"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/posts\/"}],"about":[{"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/types\/post\/"}],"author":[{"embeddable":true,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/users\/1\/"}],"replies":[{"embeddable":true,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/comments\/?post=19047"}],"version-history":[{"count":1,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/posts\/19047\/revisions\/"}],"predecessor-version":[{"id":19049,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/posts\/19047\/revisions\/19049\/"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/media\/19018\/"}],"wp:attachment":[{"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/media\/?parent=19047"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/categories\/?post=19047"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rtmedical.com.br\/en\/wp-json\/wp\/v2\/tags\/?post=19047"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}