
AI performs as well as pathologists in scoring lung cancer treatment response
Key Takeaways
- Neoadjuvant nivolumab plus chemotherapy adoption has elevated %RVT as an early post-resection decision tool, but manual IASLC scoring is labor-intensive and interobserver-variable.
- AIM-TC produced robust MPR classification across four cohorts, delivering 91% accuracy, 90% sensitivity, 93% specificity, and AUC ≥0.95 in each cohort.
AI pathology tool measures lung cancer treatment response after neoadjuvant chemoimmunotherapy with 91% accuracy, speeding NSCLC care while needing review for complete response.
An artificial intelligence (AI) algorithm scored a pathologic response to neoadjuvant chemoimmunotherapy in non-small cell lung cancer (NSCLC) with 91% accuracy compared to expert pathologists across four international patient cohorts, according to a study published Sept. 5 in
Pathologic response, also known as the percentage of a residual viable tumor (%RVT), measures how a lung tumor responded to chemotherapy with immunotherapy given before surgery. As neoadjuvant chemoimmunotherapy has become standard for early-stage, resectable NSCLC, %RVT has emerged as an early indicator for follow-up care, the study said. However, scoring tissue by eye takes time, and results can vary widely from one pathologist to the next. This can make it challenging to use as a reliable, consistent way to measure how well the treatment worked.
That shift in care took place after the
Daniel Curto, M.D., and Susana Hernandez, M.D., co-authors from the pathology department at Hospital Universitario 12 de Octubre in Madrid, led the study with their team. After using AI in earlier work to estimate lung tumor content, the researchers wanted to test whether an AI-based %RVT score would match expert pathologist response and hold up across different hospitals, scanners and patient populations.
How the study was designed
The study collected data from four retrospective NSCLC cohorts of patients treated with neoadjuvant chemoimmunotherapy and surgery. Thirty-eight patients from Hospital Universitario 12 de Octubre in Madrid (January 2020-August 2024), 49 from Peking University Cancer Hospital in China, 32 from a multicenter cohort in Germany and 16 from Spain's NADIM clinical trial were all included. In total, 135 patients and 1,344 hematoxylin and eosin-stained tumor slides were evaluated.
Pathologists scored %RVT manually under the International Association for the Study of Lung Cancer's (IASLC) response criteria, and a weighted consensus of those scores served as the reference standard. An AI algorithm called AI-based tumor cellularity measurement (AIM-TC), built by PathAI, scored the same slides independently. To check how closely the AI matched the pathologists, researchers used measures including the intraclass correlation coefficient to compare %RVT scores directly, Cohen's kappa to compare whether cases were called a major response (10% or less residual tumor) or not and standard accuracy measures to see how well the AI's calls held up.
What was found
The AI's major pathologic response (MPR) calls matched the reference standard closely with a combined accuracy of 91%, sensitivity of 90% and specificity of 93%, followed by an area under the curve of at least 0.95 in every cohort. Agreement on %RVT itself, measured by intraclass correlation, ranged from 0.77 in the China cohort to 0.88 in the Germany cohort. Manual and AI-based scoring disagreed on MPR classification in 12 of 135 cases, or 8.9%. The number of tumor cells on a slide was also linked to %RVT under both methods, though only modestly, with Kendall's tau-b values of 0.48 for manual scoring and 0.53 for AI-based scoring.
The authors also flagged a gap that was expected: the AI never called a single case a complete pathologic response (pCR), meaning %RVT of zero, though pathologists did in a small number of cases. They said this deserves a closer look, since pCR is the response category most closely tied to long-term outcomes, and suggested that general-purpose AI models like this one simply aren't trained on enough samples with almost no tumor left.
“In practical terms, this model cannot be used autonomously to adjudicate pCR,” the authors wrote, adding that any real-world use should set a low-%RVT range where a pathologist must review the case.
They pictured slides being scanned for AI review as a routine step, with a pathologist checking the results before they're used in care.
Strengths, limits and what's next
The study's main strength is its scale of our cohorts scanned on three digital pathology platforms across Spain, China and Germany, compared against a standard reference response built to IASLC recommendations, with a third pathologist reviewing every case.
However, follow-up was too short and cohorts too small. A secondary analysis of lymph node specimens, tested only in the 11-patient internal cohort, showed poor agreement between manual and AI-based scoring, with a Cohen's kappa of 0.08, driven largely by the algorithm misreading germinal center cells and artifact-damaged lymphoid tissue.
On lymph nodes specifically, the authors said the findings “should be regarded as preliminary and hypothesis-generating,” and that while the small sample size and wide confidence intervals “preclude a definitive characterization of its nodal performance,” the algorithm isn't yet fit for that purpose.
The authors concluded that dedicated model development and independent validation will be required before AI-based lymph node scoring could reach the same levels as the primary tumor results.
Related to this article










