The question most hospitals ask before buying artificial intelligence for radiology is “which algorithm?” According to radiologist Anand Singh, speaking to Radiology Business, that is the wrong question — or at least the third in line. Two come first: what problem are we trying to solve, and how will we measure success and value once the system is in production. Singh argues institutions should adopt the concept of clinical AI stewardship and treat AI as a continuous life cycle rather than a purchase with an end date.

Measuring turnaround is not measuring value
Turnaround time and productivity are the metrics in every business case, and Singh acknowledges they matter. The problem is that they answer only a slice of the question. His list of what should be on the ledger: installed capacity, patient access, safety, consistency across readers, and downstream effects on cost, revenue and clinical outcomes.
The distinction is practical. A triage algorithm that shaves 15 minutes off an exam already reported in two hours changed nothing for the patient. The same algorithm applied to a stroke queue, where minutes decide thrombolysis eligibility, changes everything. Same product, same technical gain, incomparable clinical value. Deciding where the gain matters is a management call, not a vendor call.
Platform or collection of point solutions
The second strategic question is how multiple applications will coexist. Singh argues a platform-based approach reduces interoperability friction and simplifies monitoring and governance, while a proliferation of independent tools creates a management debt that grows unnoticed.
The arithmetic is uncomfortable. A health system with dozens of AI applications has to monitor each one, understand each one’s failure modes and sustain separate workflows for each. Multiply that by revalidation at every version bump, by contract, by integration point with PACS and RIS. A common architecture — detection, reporting, quality and safety under one roof — does not eliminate the work, but it makes it finite.
Integration is the other side of the same coin: AI that demands an extra login, a parallel alert or one more workflow step simply does not get used. A radiologist under volume pressure routes around anything that costs seconds per exam. That is not technophobia; it is shift arithmetic. Tools born inside the reading workstation, as when NewVue added native reporting to the radiologist cockpit, start with a structural advantage over anything that opens a new window.
Drift: when a cleared algorithm quietly stops working
Here is the most important technical argument in the interview, and the least discussed in procurement rooms. Singh notes that changes in imaging protocols or acquisition parameters can affect algorithm performance. The concept deserves unpacking.
A model learns the distribution of the data it was trained on. When that distribution moves, performance falls — what the literature calls dataset shift. The most common variant in radiology is covariate shift: a scanner swap, a different reconstruction kernel, adoption of iterative or deep-learning reconstruction, a dose adjustment, a change in contrast phase, a firmware update. None of that changes the diagnosis; all of it changes the pixels. There is also concept drift, when the relationship between image and outcome itself shifts — a new patient population, a new diagnostic criterion.
The dangerous trait is silence. The model throws no error, does not crash, does not warn. It keeps returning probabilities with the same confident face and simply gets it right less often. Without continuous measurement, the decline surfaces only as an isolated adverse event nobody connects to the algorithm. One cheap and underused canary: track the algorithm’s positive rate over time. If the share of flagged exams moves while the patient mix has not, something shifted — and it is worth investigating before the clinical route finds it for you.
The unit of analysis is the human + AI pair
Traditional monitoring asks whether the model still delivers the promised accuracy. Singh argues the question must be broader: how is the performance of human and machine together. It is not enough to know the model is right; you need to know how the clinician interacts with its output and whether the technology shifted clinical decisions in unplanned ways.
Two well-documented effects justify the concern. Automation bias is the tendency to accept the machine’s suggestion without the scrutiny applied to a colleague — and it grows as the algorithm gets better, because trust calibrates to accumulated experience. Its mirror is alert desensitisation: an excess of false positives trains the reader to ignore the flag, including when it is right. In both cases the model’s standalone metric stays intact while the sociotechnical system degrades. That decoupling is exactly what showed up when Radiology Partners pressed the FDA for clearer rules on imaging AI, pointing at the gap between authorisation and real-world behaviour.
Clinical infrastructure, not just IT infrastructure
Singh insists AI cannot be treated as a purely IT matter of compute, storage, cloud and cybersecurity. What is missing is clinical infrastructure: named people to give feedback, review outputs, monitor safety signals and manage the interaction between human and algorithm. In practice that means local validation before go-live, monitoring after it, named clinical oversight, staff training and a formal channel to report suspected failure.
The external framework is moving the same way. The FDA now accepts predetermined change control plans, which let a manufacturer update a model within pre-approved limits — shifting onto the provider the burden of noticing that the version changed. The ACR runs registry and post-deployment monitoring programmes for precisely this purpose. And the European AI regulation classifies systems embedded in medical devices as high risk, with post-market monitoring obligations phasing in through 2027. The message converges: authorisation became the start of the process, not the end.
Three questions before signing
Singh boils the decision down to three questions that fit on one page: what problem are we solving; how will we measure success and value; and how will we scale responsibly while managing both the technology and the human-AI interaction. A hospital that does not answer all three before signing will answer them afterwards, under worse conditions.
Outside the US the picture adds a layer. In Brazil, software as a medical device is regulated by Anvisa while general AI legislation is still moving through Congress, leaving buyers without a consolidated national post-market yardstick. Supply, meanwhile, keeps growing: radiology already accounts for 76% of all FDA AI device authorisations, and much of that catalogue arrives through local distributors. Without validation on local data — a different population, equipment fleet and protocol set from the training environment — the vendor’s metric is a promise, not a measurement.
The lowest-regret path is to start small and measurable: one defined clinical problem, a baseline recorded before go-live, a named clinical owner and a date set to reassess. If the value does not show up on the dashboard, switching it off is a legitimate decision — and a cheap one, if the contract anticipated it.
Source: Radiology Business

