Skip to main content

Same images, different reports

A study published in the Journal of the American College of Radiology (JACR) put a number on something most imaging directors suspect but rarely measure: false-positive rates in low-dose CT lung cancer screening range from 11% to 22% across radiologists. In randomly drawn pairs of readers, one of the two was on average twice as likely to flag a suspicious finding that later proved harmless. Same images. Same clinical question. What changes is who reads the case.

Radiologist reviewing a chest CT on a PACS workstation with axial, sagittal and coronal lung reconstructions
In lung screening, each reader’s personal threshold weighs as much as the acquisition protocol

What the JACR study measured

The researchers worked from a large base: 1,400 radiologists, each with at least 50 low-dose CT (LDCT) screening exams, for roughly 200,000 exams in total. That 50-read floor is not a throwaway methodological detail. It removes the occasional reader, whose rate would swing on sampling luck alone, and keeps only clinicians with enough volume for the observed rate to reflect genuine reading behavior.

From there, the team calculated false-positive rates separately for baseline exams (each participant’s first scan) and for follow-up exams. The split makes clinical sense: at baseline everything is new and every nodule has to be characterized, while at follow-up the reader has a temporal comparison and can judge stability. The spread showed up in both settings anyway.

The next finding is what gives the paper its weight. Radiologists with above-median false-positive rates also had higher sensitivity: 97% versus 92% for colleagues below the median. The “alarmist” reader was not simply making mistakes — that reader was catching more cancer. This shifts the conversation from competence to calibration. There are no good and bad readers here, only different thresholds of suspicion.

The 100-alarms-per-cancer paradox

Here is the number that belongs in every screening committee meeting: among baseline exams read by radiologists with below-median false-positive rates, there would be roughly 100 fewer false-positive screens for each true-positive cancer missed. One hundred avoided alarms at the cost of one missed diagnosis.

CT scanner in an exam room prepared for a low-dose chest computed tomography study
Acquisition technology matured; standardizing interpretation lagged behind

Which scenario is preferable? The answer is not technical, it is about values — and it likely depends on who is answering. For the participant who gets a recall, undergoes another scan, waits three months and learns it was scarring, the cost is anxiety, time and extra exposure. For the participant whose nodule slipped through, the cost may be a life. Most patients, put in front of that choice, will accept the false alarm. Health systems, which pay for the downstream cascade of exams and procedures, tend to reason differently.

Worth remembering: a false positive in lung screening rarely means an immediate biopsy. It usually means reclassification to Lung-RADS 3 or 4A, a follow-up CT in three or six months and, in a fraction of cases, PET/CT or needle sampling. The harm is real but diffuse — a little anxiety, a little cumulative dose, a little occupied scanner time. Exactly the kind of cost that accumulates without ever showing up in a dashboard.

Why mammography faced this first

Lung screening participation is climbing — from single digits a decade ago to 19% in a recent survey — but remains far from the roughly 75% seen in breast cancer screening. That maturity gap explains a lot. Mammography already went through the inter-reader variability problem and built a response: interpretation standardization programs, automated feedback systems that return individual metrics to each reader, targeted educational interventions and selective double reading for the most uncertain cases.

The JACR authors propose exactly that path for LDCT. The recommendation looks modest and is demanding in practice, because it means measuring each reader’s individual performance and handing that number back to them — something services usually avoid to keep the peace. The mammography experience shows it works, and also shows it only works when the data returns as a calibration tool rather than a punitive ranking.

There is a direct parallel with efforts to standardize image metrics elsewhere in CT. The proposal of a five-star scale for rating CT image quality comes from the same diagnosis: where there is no shared scale, opinion fills the gap. And note that the technology already did its part — average CT radiation dose fell about 22% over a decade in the United States, driven by iterative reconstruction and automatic tube current modulation. The bottleneck moved from the scanner to the interpretation.

What imaging services can do now

For anyone running a screening program, three measures require no new capital. First, measure: pull the Lung-RADS distribution per reader from the RIS or PACS and compute each one’s recall rate. Without that number, any discussion of calibration is anecdotal. Second, benchmark externally — the study’s 11% to 22% band gives a plausible interval to locate your own service. Third, institute targeted rather than universal double reading: jointly review only cases classified as 3 and 4A, where disagreement concentrates.

There is also an efficiency gain here that usually goes unnoticed. Every avoided recall is a freed scanner slot and one less report in the queue — and the queue is contemporary radiology’s structural problem, as shown by data indicating that radiology report turnaround time grew 177% over a decade in the United States. Calibrating thresholds of suspicion is not only clinical quality; it is capacity management.

Limits and what comes next

The study is retrospective and observational, and nothing in it proves that standardizing interpretation improves mortality outcomes. Nor does it separate how much variability comes from the reader and how much from the context they work in: acquisition protocol, equipment quality, availability of priors for comparison, productivity pressure. A radiologist without a prior study in the PACS is structurally more likely to call a nodule suspicious — and that is an interoperability problem, not a calibration one.

The predictable next frontier is AI as a triage second opinion, promising to narrow inter-reader spread by offering a quantitative reference for nodule volumetry and growth. The promise is sound and the evidence is still young. While it matures, the available instrument remains medicine’s oldest: measure what you do, compare with peers, adjust. The JACR study does not resolve the tension between detecting more and alarming less — it merely puts a price on it. One hundred to one.

Source: The Imaging Wire