Skip to main content

A Five-Star Scale for Judging CT Image Quality

Researchers from the International Atomic Energy Agency (IAEA) and Massachusetts General Hospital (MGH) have proposed a deliberately simple metric for CT image quality in the European Journal of Radiology: a five-star rating, in the same spirit as the scores anyone has given a restaurant or an app. The idea sounds light, but it targets a real imbalance in radiology. Radiation dose has numbers, units and monitoring software. Image quality, until now, had no universal scale — it had adjectives.

CT scanner room with five gold stars overlaid, representing the new five-star scale for CT image quality
The new scale rates a CT image against the dose it took to produce it

That imbalance has practical consequences. Protocol committees can argue about CTDIvol and DLP to two decimal places, but when the question becomes “did the images look good?”, the conversation slides into personal preference. The pattern is familiar in any department: the reader who likes smooth images pushes the protocol upward, and nobody has a shared instrument to say that the extra sharpness was not worth the extra dose.

Why Two Stars Is Worse Than Five

Here is the inversion that makes the proposal interesting. On almost every quality scale, sharper is always better. Not on this one. The system builds the dose-noise dilemma into the rating and explicitly penalizes the over-pretty image — the one acquired with dose the diagnosis never required. The scale reads as follows:

  • Five stars — acceptable, despite high image noise associated with low-dose CT.
  • Four stars — acceptable for all parts within the region of interest.
  • Three stars — acceptable, despite unintended high noise and/or some artifacts.
  • Two stars — excellent quality, with unjustifiably low image noise and no artifacts.
  • One star — non-diagnostic, due to excessive noise or artifacts, or poor contrast.

Read the two-star line again. It describes the image many departments display with pride and files it as nearly the worst possible outcome — not because it looks bad, but because the beauty was paid for with dose the patient did not need. It is a subtle value inversion and, for anyone working in optimization, a liberating one: at last there is a grade that separates “excellent image” from “adequate image”.

Six Hospitals, Five Countries, 2,700 Patients

To test whether the scale survives outside the paper, the IAEA and MGH teams ran a study across six hospitals in five European countries. CT scans from roughly 2,700 adults were assessed, and three radiologists per hospital were trained on the scale before scoring images. The design was deliberate: the point was to measure how much of the score depends on who is looking.

Reader concordance reached 93% at some sites, and overall discordance was rare, at 1.6% — a low figure for a subjective metric applied by different observers, in different departments, with different scanners and protocols. More relevant to medical physics: two-star ratings were associated with higher radiation dose across nearly all patient weight groups. The scale does not merely describe perception; it can flag, straight from the clinical read, where dose is being wasted.

It helps to understand why dose-noise trade-offs make this kind of judgement so treacherous. In CT, the standard deviation of quantum noise falls only with the square root of dose:

$$\sigma \propto \frac{1}{\sqrt{D}}$$

where $\sigma$ is image noise and $D$ is dose. The consequence is unforgiving: halving noise requires quadrupling dose. Every extra step of visual polish costs disproportionately in radiation — and that invisible cost is exactly what the two-star grade makes explicit in a quality report.

What Changes in Day-to-Day Practice

For the people running the scanners and owning the quality assurance program, the usefulness is immediate. A short scale that can be memorized and taught in minutes can be applied in periodic image audits, cross-referenced with the scanner’s own dose records, and used as a protocol-review trigger. When a batch of abdominal exams starts collecting two stars, there is finally an objective argument to cut mAs or lean harder on iterative reconstruction, rather than a debate about taste.

The instrument fits neatly into the broader optimization trend already visible in the data: in the United States, average CT dose fell about 22% over a decade, driven by iterative reconstruction, automatic tube current modulation and systematic protocol review. The star scale supplies the missing counterpart: a way to check whether that dose reduction cost diagnostic capability or not.

The effect is sharper still in screening programs, where low-dose CT is the rule and noise is accepted by design. Noisy exams can be perfectly diagnostic for the question that matters. On a traditional scale, those exams would score badly. On this one, they are five stars — and that change of reference point may relieve the informal pressure to push dose upward in population programs, a concern that also surfaces in international efforts to widen access to cancer imaging coordinated by the IAEA.

Limits, and What Comes Next

There are honest limits. The scale is subjective by construction and does not replace physical metrics such as contrast-to-noise ratio, noise power spectrum or task transfer function, which remain the backbone of instrumental quality control. The study used three readers per site, adult patients and European protocols; validation in pediatrics, in high-volume practices and in emergency settings is still pending. And there is the predictable risk that the grade, once written into a contract or a performance indicator, gets gamed like any other metric.

Workflow integration will decide the outcome. If the rating can be entered with one click inside the PACS, next to the report, it survives. If it requires a parallel spreadsheet, it dies within three months. Departments already tracking productivity pressure — the same ones watching report turnaround times climb 177% over a decade — will not adopt anything that adds clicks without returning insight.

The open question the authors leave behind is whether the system becomes an integral part of radiology’s dose-management effort or remains an elegant thought experiment. For the medical physicist and the technical lead, though, the conceptual gain is already banked: noise has stopped being a defect and started being evidence that dose was matched to the clinical question.

Source: The Imaging Wire