Skip to main content

What radiologists expected from AI, and what they report

Half of breast radiologists already use an FDA-cleared artificial intelligence tool for cancer detection — and in their own assessment, the gain came in well below what they imagined before installing the software. That is the finding of a survey of 215 Society of Breast Imaging (SBI) members published in Clinical Imaging and led by Joud Almogati and colleagues at UC San Diego Health. The paper does not measure algorithm accuracy. It measures the gap between expectation and lived experience, and that gap is wide.

Patient undergoing digital mammography assisted by a technologist in a breast imaging service
High adoption, modest perceived impact: the state of mammography AI in 2026

What the SBI survey measured

The sample consists of SBI-affiliated breast radiologists, and adoption broke down as follows: 47% have already implemented breast cancer detection AI, 11.2% plan to implement it, and 41.8% use no diagnostic AI at all. In practice, the specialty splits roughly down the middle between those living with an algorithm and those watching from outside.

The design is declarative — these are readers’ subjective perceptions, not outcome measurements pulled from a registry. That is a limitation and also the point. No purchasing committee decides based on sensitivity in an external validation set; it decides based on what its own radiologists say they are feeling after six months of use. That layer rarely gets published.

Recall, biopsy and burnout: three mismatches

The core numbers compare what participants expected before implementation with what they report observing afterward. Worth reading slowly:

  • Lower recall rates — 59% expected a reduction; 35% report having seen one.
  • Fewer unnecessary biopsies — 36% expected it; only 9% report observing it.
  • Less burnout — 56% expected relief; 29% report actual relief.

The steepest drop is in biopsy: from 36% to 9%, a four-to-one ratio between hope and experience. It makes sense in workflow terms. AI acts at the moment of reading, flagging suspicious regions; the decision to biopsy depends on additional characterization, targeted ultrasound, clinical correlation and often patient preference. A probability marker on a mammogram is simply not the bottleneck in that decision chain.

The burnout item deserves particular attention from anyone managing a team. The expectation that AI would ease cognitive load was held by a majority, and reported relief landed at roughly half that. European surveys point the same way: about 70% of radiologists say they perceive no reduction in clinical workload with AI, and some worry the tool adds tasks — validating marks, documenting disagreement, explaining the output to the patient.

Why AI became a second opinion, not a decision

The reported usage pattern is telling: most see AI primarily as a second opinion, and few treat it as a deciding factor in management. According to Almogati, only a minority treat the algorithm’s output as decisive.

There is a solid technical reason for that stance, and it is not conservatism. A suspicion score between 0 and 100 does not itself carry the positive predictive value applicable to the specific patient in front of you — that depends on local prevalence, breast density, family history and whether priors are available for comparison. Without local calibration, the number is a signal, not a verdict. The radiologist treating the score as a suggestion is doing exactly what statistics recommend.

Then there is the threshold problem. Detection tools operate at a cutoff that trades sensitivity against specificity. Moving the cutoff to avoid missing cancer increases marks to work up; moving it to reduce false alarms reintroduces the risk of a missed lesion. No setting improves both ends simultaneously — which is why the promise of “less recall and more detection” at the same time tends to dissolve in routine practice.

Cost and institutional support: the barriers that remain

Among those who have not adopted, the two dominant barriers are not clinical: cost and lack of institutional support. That shifts the debate from accuracy to financing. A mammography detection algorithm is a recurring per-exam or subscription expense, without a dedicated reimbursement code in most settings, and the benefit shows up diffusely — in a shorter queue, an avoided recall, a sense of safety.

The practical reading for any service is direct. Departments weighing breast AI need to decide beforehand which indicator will justify the invoice: recall rate, turnaround time, detection per thousand exams, or the ability to absorb volume without new hires. Without a metric chosen before the contract is signed, the predictable result is the one captured in this survey — tool installed, used as a second opinion, benefit hard to demonstrate.

None of this means the technology is standing still. The recent FDA clearance of DeepHealth’s AI for breast ultrasound shows the frontier moving beyond mammography. And the question of which modality best complements screening remains open, as in the debate over the role of ultrasound in screening in the tomosynthesis era.

How to use these findings in practice

Three practical recommendations follow from the work. First, align expectations in writing before deployment. If the team expects lower recall while the vendor promises a sensitivity gain, the project starts misaligned and ends in frustration — what the authors frame as a prerequisite for sustainable integration.

Second, measure the baseline before switching the algorithm on. Recall rate, benign biopsy rate and detection per thousand exams must exist as your own historical numbers; without them, any later evaluation is impression.

Third, budget for the communication cost. Once AI participates in the read, patients start asking what the computer thought. The ACR has already issued guidance on AI-generated report summaries for patients, a sign that the explanation layer is now part of the job rather than an extra.

Study limits and outlook

The survey carries honest limitations: 215 respondents from a subspecialty society do not represent global breast radiology, there is self-selection bias among those who answer an AI questionnaire, and every impact measure is self-reported without registry verification. It is entirely possible that AI is reducing recall in services whose radiologists do not perceive the difference — perception and effect are not the same thing.

Even so, the message is useful precisely because it is uncomfortable. The technology entered routine practice in half of specialized services without delivering the relief that was advertised. The next evidence cycle needs to measure auditable outcomes rather than expectations: verified recall from registry data, benign biopsies per thousand exams, actual reading time. Anyone buying breast AI in 2026 without defining those indicators will, three years from now, simply reproduce the survey that was just published.

Source: Clinical Imaging