17 SEP 2026 — An artificial-intelligence model trained on 2,396 lung-cancer patients outperformed every biomarker oncologists currently use to decide who will benefit from immunotherapy. Moved to an external cohort, its scores fell to between 0.55 and 0.72.
Both findings are in the same paper, published in Nature Medicine on 14 September. Only the first reached most of the coverage.
What the model predicts
I3LUNG addresses a narrow, expensive question. Patients with stage IIIC to IVB non-small-cell lung cancer may or may not respond to immunotherapy, and the tools used to guess in advance are weak. PD-L1 expression is the standard, and its limitations are well known to the doctors relying on it.
The researchers built models from routinely collected data. Clinical records and blood results came first; computed tomography imaging, digital pathology and genomics were layered on top. The work enrolled patients across six centres in Italy, Germany, Greece, Israel, Spain and the United States, and was funded by the European Union's Horizon 2020 programme.
The number the paper defends
In the independent test set, models built from clinical and blood data alone reached an area under the curve of up to 0.77. The paper reports they significantly surpassed PD-L1, ECOG performance status, the neutrophil-to-lymphocyte ratio, lactate dehydrogenase and the Lung Immune Prognostic Index.
That is the substantive claim. Beating five established predictors on an independent test set using data a hospital already collects is a useful result, and it does not require the imaging and pathology layers at all.
The number that fell
The result thins in external validation. The paper gives an AUC range of 0.55 to 0.72 there and attributes the drop to population differences. The authors state plainly that the multimodal gains were not consistently reproduced in independent testing or external validation.
An AUC of 0.55 is close to a coin toss. The range matters more than its midpoint, because it says the model's usefulness depends on which hospital's patients it meets.
The institutional announcement led with 0.88, the figure for the full multimodal model. That number is in the paper, but it is not the number the paper defends once the model leaves the population it learned on. A reader who saw only the release would not know the 0.55 to 0.72 range existed.
What the doctors did with it
The more interesting half of the study put the tool in front of physicians. Asked to predict disease control, their sensitivity rose from 0.72 without the explainable-AI tool to 0.87 with it. Non-experts improved most, which is what a decision-support tool is for.
Overall accuracy moved from 0.57 to 0.65. Hold onto that. A large gain in sensitivity with a modest gain in accuracy means clinicians are readier to say a patient will respond, but they are not correct nearly as often as the sensitivity figure suggests on its own.
Some coverage has reported the 0.72 to 0.87 movement as an AUC. It is a sensitivity score for one prediction task. The distinction matters: the overall accuracy gain was eight points, not fifteen.
Six centres, none in Asia
The six participating centres are in Europe, Israel and the United States. For readers here that is the whole question, because the paper's own finding is that performance moves with the population.
Lung cancer in Southeast and East Asia is not the same clinical picture. The proportion of never-smokers is higher, EGFR-mutant disease is substantially more common, and the treatment sequence around immunotherapy differs accordingly. A model whose external validation already wobbles inside Europe has not been tested against any of that, and nothing in the paper claims otherwise.
This echoes our earlier clinical AI reporting, including the ArteraAI digital-pathology clearance. The training population is a property of the model, not a footnote to it.
What to watch
Start with the prospective phase. The decision-support system is now being validated prospectively in more than 2,000 patients, and a prospective result is a different class of evidence from a retrospective one. If the clinical-and-blood model holds near 0.77 there, the tool has stronger footing.
Then the multimodal question. Imaging and pathology are expensive to standardise across hospitals, and the paper's own caveat is that their benefit did not reproduce consistently. If the cheap model survives validation and the expensive one does not, that is a finding about what to deploy, not a disappointment.
Then who is asked next. Generalisation is the stated weakness, so the test that matters for this region is an Asian cohort. Until someone runs it, the honest position is that this tool has been shown to work on populations that are not ours.