An FDA-cleared aneurysm algorithm flagged 55 genuine intracranial aneurysms that radiologists had not reported across 3,856 head CT angiograms — a 39% relative lift in detection. The figure comes from a prospective Northwell Health study published 15 September in the Journal of the American College of Radiology. It also comes with the half of the ledger that headlines tend to drop: the same tool raised 46 false alarms and missed 30 aneurysms the radiologists caught.

What the study actually counted
The design is what gives the result its weight. Investigators ran the Aidoc algorithm in shadow mode, meaning it processed studies alongside the clinical workflow while its output stayed invisible and changed no care decision for the duration. Radiologists read all 3,856 angiograms as they normally would, blinded to the tool. Where the two disagreed, independent expert neuroradiologists went back to the images and settled which read was right.
Across the whole cohort, reader and machine landed on the same answer in more than 96% of examinations. Everything interesting sits in the remaining 4%. The algorithm raised 101 findings on its own that the report did not contain; 55 turned out to be aneurysms and 46 were false positives. Going the other way, radiologists reported 30 confirmed aneurysms the algorithm never marked. Incremental true detections outnumbered the false alerts, which the authors frame as a favourable benefit-to-burden ratio.
In conventional terms the tool was the more sensitive reader (84.6% against 71.8%), while the radiologist held the better positive predictive value (92.7% against 78.2%). Specificity and negative predictive value came out broadly similar. “Intracranial aneurysms can remain clinically silent until rupture, which can result in devastating consequences for patients,” said senior author Pina C. Sanelli, MD, MPH, FACR, of the Zucker School of Medicine at Hofstra/Northwell, who directs the Neiman Health Policy Institute’s PRIME Center.
Sensitivity and precision describe opposite mistakes
Worth pulling the two metrics apart, because they answer different questions. Sensitivity asks: of every patient who truly had an aneurysm, what share did the reader flag? Positive predictive value asks: of every flag the reader raised, what share was real? One measures what slips through, the other measures how much a single alert is worth.
Put it in shift terms. For every 100 marks the algorithm made on its own, roughly 78 corresponded to an actual aneurysm. For every 100 aneurysms a radiologist called, roughly 93 held up. The human is more conservative and therefore more credible when speaking; the machine is more aggressive and therefore finds more while crying wolf more often. No threshold setting lifts both numbers at once on the same detector — lowering the alert threshold buys sensitivity and spends precision, raising it does the reverse. Picking that operating point is a clinical judgement, not an engineering one.
One statistical footnote explains why specificity and NPV came out so close: the overwhelming majority of the 3,856 studies were negative. At low prevalence nearly any detector nails the negatives and both metrics pin near the ceiling, which is exactly why quoting them alone hides the real gap between readers. The same shape showed up when the vendor’s pulmonary embolism tool was measured in the field — see our earlier look at AI for PE detection on CTPA in real-world use.
Where the extra aneurysms came from
This is the part the press release underplays. Sanelli noted that many of the additional aneurysms the tool surfaced “were among the smallest lesions” — a sentence that reframes the whole result, because size is the single variable that most shapes the natural history of an unruptured aneurysm.
Unruptured intracranial aneurysms are common. The reference meta-analysis by Vlak and colleagues in Lancet Neurology (2011) put prevalence at roughly 3.2% among adults without comorbidity, and the vast majority never bleed. That is precisely why the decision to treat is not made on the finding itself but through risk scores: PHASES (population, hypertension, age, size, earlier subarachnoid haemorrhage from another aneurysm, site) estimates five-year rupture risk; the UIATS weighs arguments for intervention against arguments for observation; ELAPSS projects growth. Small lesions at low-risk sites in patients without aggravating factors usually land in surveillance, not in a clipping or coiling suite.
So those 55 extra findings are real and not trivial, but most of them convert into imaging follow-up and risk-factor control — blood pressure, smoking cessation — rather than into 55 haemorrhages averted. Add that the scores themselves discriminate imperfectly, with recent series reporting that PHASES and UIATS separate ruptured from unruptured aneurysms less cleanly than clinicians would like, and the diagnostic gain lands squarely in territory where management is still argued case by case.
Care setting decides whether the alert pays for itself
The most operationally useful result is not in the headline percentages. Performance swung sharply by where the scan was ordered. Among inpatients, the algorithm added 18 aneurysms while producing only 7 false-positive alerts — its strongest showing on every metric. The emergency department also came out favourable. In the outpatient setting it contributed just four extra detections and generated more false positives than true ones.
“A likely explanation is that higher-acuity inpatient and emergency settings involve more clinically complex examinations,” said Matthew Barish, MD, professor of radiology at the same school and CMIO of clinical shared services at Northwell. The practical reading is blunt: switching the algorithm on across every exam type dilutes the benefit and imports the noise. Deployment scope is a configuration decision with clinical consequences.
There is a downstream problem too. A small aneurysm found at 3 a.m. raises questions most departments have not written down: who tells the patient, who books the follow-up, at what interval, and on which modality — non-contrast MR angiography or a repeat CTA with its extra radiation and iodine load. Without that pathway, a 39% detection lift translates into a matching lift in anxiety and outpatient return visits. False alerts cost something as well, in confirmatory imaging and in reader trust, as we saw with an AI-flagged pneumothorax that was never on the radiograph.
What shadow mode cannot tell you
The central limitation is baked into the method. Shadow mode measures what the algorithm would have found, not what happens once its output reaches the reading station. Nothing here speaks to automation bias, to added reading time, or to how often a radiologist would overrule a correct flag. Nor are there patient outcomes: rupture, treatment and mortality were not tracked.
Scope limits apply as well — one integrated health system, one vendor, one model version. That is the exact point raised by Elizabeth Rula, PhD, executive director of the Neiman Health Policy Institute and a study co-author, who argued that organisations should judge AI by how far it improves physician performance in real use rather than by the numbers achieved in its original validation environment, making post-deployment monitoring part of the job instead of an optional extra. That expectation only hardens as the approved catalogue grows: radiology already accounts for 76% of all FDA AI device authorizations, and Aidoc itself keeps stacking milestones, including the $150 million round led by Goldman Sachs and Nvidia. Clearance is the entry ticket; measuring in the department is the work that remains.
Source: ITN Online




