user Admin_Adham
4th Sep, 2026 12:00 AM
Test

AI Speeds Up Chest Radiograph Readings but Lowers Accuracy

TOPLINE

In a prospective real-world crossover study, the use of four commercial AI tools reduced interpretation times and boosted reader confidence but failed to improve chest radiograph accuracy, with 71% of AI-prompted revisions converting correct judgements to incorrect ones and increased false positives.

METHODOLOGY

  • Researchers conducted a prospective crossover reader study involving 1200 consecutive patients undergoing chest radiography (1861 total radiographs) in emergency, inpatient, and outpatient settings at a single centre in Germany between April and June 2025.
  • Five radiology residents with 1-6 years of experience independently assessed each case under five conditions: without AI and with each of four commercially available AI algorithms, with a 14-day washout period and randomised case order between sessions.
  • Readers evaluated each case for five pathologic findings (pulmonary infiltrates, pleural effusions, mediastinal masses, pneumothorax, and pulmonary nodules).
  • The primary outcome was diagnostic performance assessed using sensitivity, specificity, positive predictive value, negative predictive value, accuracy, and area under the receiver operating characteristic curve, with the reference standard being the final clinical radiology report with confirmatory CT when available.
  • Secondary outcomes included reading time per case, diagnostic confidence measured on a 5-point Likert scale, escalation to senior review, and CT recommendations.

TAKEAWAY

  • None of the four AI algorithms improved diagnostic accuracy for any of the five findings; for pulmonary nodules, the accuracy declined from 98.00% without AI to 93.50% with algorithm 1 (P = .004), 91.00% with algorithm 2 (P < .001), 93.00% with algorithm 3 (P = .002), and 91.50% with algorithm 4 (P < .001) in one reader.
  • For pleural effusion, the accuracy of another reader decreased from 97.00% without AI to 94.00% with algorithms 1 and 2 (P = .031 for each) and 93.50% with algorithm 3 (P = .016); however, no significant decrease occurred with algorithm 4 (P = 1.000). 
  • Across 24,000 reader-finding decisions under AI assistance, readers changed their call in 525 instances, of which 71% converted a correct unassisted judgement to incorrect, whereas only 29% corrected an initially incorrect judgement (P < .001), yielding a harmful-to-beneficial ratio of 2.45:1. 
  • Of five readers, three achieved significantly shorter interpretation times with AI assistance (P ≤ .031 for all comparisons) and four reported significantly higher diagnostic confidence (P ≤ .003 for all comparisons); one reader had a significant reduction in senior consultation rates with all AI tools (from 6.5% without AI to 0.5% with algorithms 1 and 3, 1.0% with algorithm 4, and 0.0% with algorithm 2). 

IN PRACTICE

"While this study prospectively evaluated time-to-report, escalation burden, and avoidable follow-ups, the true value of commercial AI in chest radiography will depend on whether such efficiency gains translate into meaningful economic benefits at the departmental and health-system level," the authors wrote.

"[T]he algorithms evaluated here were provided under free evaluation licenses, so acquisition, integration, and licensing costs were neither incurred nor assessed," they added.

SOURCE

The study was led by Tristan Lemke, MD, and Alexander W. Marka, MD, TUM University Hospital Rechts der Isar, Munich, Germany. It was published online on August 20, 2026, in Academic Radiology.

LIMITATIONS

The study was limited by the single-centre design conducted exclusively with radiology residents in Germany, the low prevalence of mediastinal masses and pneumothoraces, the reliance on finalised clinical reports as the primary reference standard with confirmatory CT available for only 10.8% of patients, unequal reader case volumes, the reliance on binary reader responses rather than a graded level-of-suspicion scale, and potential learning and carry-over effects.

SUGGESTED FOR YOU

DISCLOSURES

The research did not receive any specific funding. One author reported receiving grants and speaker fees from the European Union (EU), the Wilhelm Sander Foundation, Canon Medical Systems Corporation, GE HealthCare, and other organisations and serving as a member of the advisory board of the EU Horizon 2020 LifeChamps Project and the EU Innovative Health Initiative Project IMAGIO.

This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.

Dive Deeper
References
Commonly Asked by HCPs


Share This Article

Comments

Leave a comment