An artificial intelligence model trained on routine chest CT scans can flag, before treatment starts, which lung cancer patients face a higher risk of immunotherapy-induced pneumonitis. Named CIPHER (Checkpoint-Inhibitor Pneumonitis Hazard EstimatoR) and developed at the University of Texas MD Anderson Cancer Center in Houston, the system reached an AUC of 0.83 in an external Johns Hopkins cohort, compared with 0.66 for a conventional radiomics model. The study was published in September in the Journal for ImmunoTherapy of Cancer.

What the study did and what the numbers show
Pneumonitis linked to immune checkpoint inhibitors (ICIs) is a potentially life-threatening lung inflammation that affects roughly 10% of lung cancer patients on immunotherapy. No test currently tells clinicians in advance who will develop it. The group led by Amgad Muneer, with Jia Wu (Imaging Physics), Ajay Sheshadri (Pulmonary Medicine) and Mehmet Altan (Thoracic Oncology) as senior authors, started from a direct question: does the CT every patient already undergoes before treatment carry enough information to estimate that risk?
To answer it, the team built CIPHER in two stages. First, the model was pretrained in a self-supervised fashion on 590,284 CT slices from 4,286 scans of 2,500 patients with non-small cell lung cancer (NSCLC), with no clinical labels at all. Second, it was adapted to an internal cohort of 347 ICI-treated patients, 33 of whom had expert-adjudicated pneumonitis. Fine-tuning used only 254 patients without pneumonitis, and performance was measured on a held-out test set of 93 patients (33 cases and 60 controls).
On that internal test, AUC ranged from 0.77 to 0.85 across five independent runs, with balanced accuracy of 83.3%, sensitivity of 80.0% and specificity of 86.7%. In external validation on 116 Johns Hopkins patients (20 cases and 96 controls), CIPHER held an AUC of 0.83 and balanced accuracy of 81.7%, correctly classifying 80 of 96 controls and 16 of 20 cases. The radiomics comparator, although sensitive (85%), misclassified more than half of the controls, with specificity of just 45.8%; the AUC gap (0.83 versus 0.66) was statistically significant (p = 0.032).
One further result matters to clinicians: after adjusting for clinical variables, each standard-deviation increase in the CIPHER score was associated with roughly a threefold higher hazard of toxicity. The model’s attention maps highlighted subpleural reticulation and ground-glass opacities, findings consistent with subtle and possibly undiagnosed interstitial lung disease (ILD). The work was funded by the NIH, CPRIT and MD Anderson, and the code is available on GitHub (WuLabMDA/CIPHER).
What a foundation model is and why it matters for CT
The term “foundation model” describes a neural network trained on a large volume of unlabeled data that learns general representations of a domain and is later specialized for concrete tasks with few examples. CIPHER combines two such techniques. The masked autoencoder hides a fraction of the patches in each CT slice and forces the transformer-based network to reconstruct what was removed; to succeed, it must internalize how normal lung parenchyma is organized. Contrastive learning, in turn, pulls together the representations of augmented views of the same scan and pushes apart those of different scans, making the descriptors more robust to variations in protocol and scanner.
That combination helps explain why the approach tends to beat classical radiomics in small cohorts. Radiomics models rely on hundreds of texture descriptors computed inside a manually segmented region and often overfit to the scanner they came from, which is consistent with the comparator’s 45.8% specificity in external validation. A network that has learned to reconstruct thousands of lungs, by contrast, carries a built-in statistical sense of what normal looks like: regions that depart from that pattern, such as juxtapleural reticular bands or faint ground-glass, are precisely the ones it reconstructs worst and therefore the ones that weigh most in the representation. The same logic lets deep learning models read cancer risk in CT scans called normal.
Immunotherapy pneumonitis versus radiation pneumonitis: why the distinction matters
For radiation oncology readers, the study is directly relevant. Radiation pneumonitis is a dose-dependent effect confined to the irradiated volume, typically appearing one to six months after treatment and following the geometry of the fields. ICI pneumonitis, by contrast, is immune-mediated, usually bilateral and diffuse, may emerge weeks or many months after the first infusion, and takes varied CT patterns, from organizing pneumonia to acute interstitial disease. In practice the two increasingly overlap: since the PACIFIC trial, consolidation durvalumab after chemoradiation has become the standard in unresectable stage III NSCLC, and in that trial pneumonitis of any grade, including radiation pneumonitis, was more frequent in the ICI arm than with placebo.
A risk score computed on the planning CT or the baseline CT could, in principle, inform decisions that today rest on clinical judgment: tightening lung dose constraints, choosing between regimens with or without immune consolidation, or stepping up imaging surveillance. The same caution applies to reirradiation, where cumulative lung dose already demands careful planning. It is worth stressing that CIPHER was not trained to distinguish the two types of pneumonitis or to predict radiation toxicity; that extension is a hypothesis, not a result of the study.
Where CIPHER would fit in the oncology workflow
The baseline chest CT is mandatory before any line of systemic treatment: it stages the disease, serves as the reference for response criteria and is often the same image used for radiotherapy planning. A model that extracts a risk score from that exam, with no additional acquisition, no extra contrast and no manual segmentation, has a low marginal cost. The bottleneck is not technical but organizational: someone has to validate the model locally, integrate it into the PACS, decide who receives the alert and what to do with it. As we argued when examining why radiology AI needs a management plan first, prognostic tools only create value when a clinical protocol is coupled to the algorithm’s output.
In Brazil and the wider region, the question is no longer theoretical. Pembrolizumab has been incorporated into the public health system (SUS) for a subgroup of metastatic NSCLC patients with high PD-L1 expression, and consolidation durvalumab is on the mandatory coverage list of the private health regulator ANS, expanding the number of patients exposed to pneumonitis risk. At the same time, ILD is underdiagnosed, and the staging CT is often the first chance to spot subtle reticulation or ground-glass, findings whose description in reports, as a recent study showed, varies widely among radiologists. An automated score could standardize that triage. For now, no device of this kind is cleared for the indication, and the regulatory landscape of imaging AI, dominated by detection tools, still includes few authorized prognostic algorithms.
Limitations and next steps
The authors themselves list the constraints. The model was trained and evaluated only in NSCLC, with a small number of events (33 internal and 20 external cases), in a retrospective design and with pneumonitis adjudicated after the fact. Four of the twenty external cases went undetected, and a false negative in this setting means a high-risk patient without enhanced surveillance. It is also unclear how the score behaves in small cell carcinoma, in previously irradiated patients or in populations whose ILD profile differs from that of the United States.
The team calls for prospective, multicenter validation before any clinical use and suggests CIPHER could guide personalized monitoring strategies and preventive interventions for the most vulnerable patients. Releasing the code makes replication easier for other centers, including those in Latin America, with CTs acquired on different scanners and protocols. The full paper is open access in JITC, and MD Anderson issued a press release on the work.
Source: AuntMinnie




