user Admin_Adham
27th Feb, 2026 12:00 AM
Test

Can Large Language Models Simplify Radiology Reports?

TOPLINE:

In a meta-analysis, large language model (LLM)-simplified radiology reports were found to be more understandable and readable for patients. Clinicians rated the LLM-simplified reports highly for accuracy and completeness; however, few of the reports contained clinically significant errors requiring human oversight.

METHODOLOGY:

  • Researchers conducted a systematic review and meta-analysis including 38 studies applying LLMs to simplify radiology reports assessed by patients, the public, or medical professionals between 2022 and 2025.
  • Included studies generated 12,922 simplified reports evaluated by 508 assessors (387 laypeople and 121 medical professionals).
  • The simplified reports covered six imaging modalities, most commonly MRI (66%) and CT (58%). GPT models (OpenAI) were used in 92% of studies, and GPT-4 was the most common among them (47%).
  • Primary outcomes were self-reported understanding of LLM-rewritten radiology reports by patients or laypersons and their quality as assessed by clinicians.
  • Secondary outcomes included objective readability metrics, including Flesch-Kincaid Grade Level (FKGL), and error rates of LLM-rewritten reports.

TAKEAWAY:

  • Patients or lay assessors rated LLM-rewritten reports as substantially more understandable than original radiologist-generated reports, with a pooled mean score of 4.04 for simplified reports vs 2.16 for original radiologist-generated reports, representing a mean difference of 2.00 (95% CI, 1.54-2.46).
  • Clinicians rated LLM-rewritten reports highly for accuracy and completeness, with a pooled mean score of 4.45 (95% CI, 4.27-4.63) and 4.53 (95% CI, 4.30-4.76), respectively; however, releasability and safety ratings were lower at a pooled mean of 3.93 (95% CI, 3.10-4.77) and 3.79 (95% CI, 3.10-4.51), respectively.
  • Readability improved substantially across all imaging modalities, with pooled mean differences in FKGL scores of -6.20 (95% CI, -6.91 to -5.48) for CT, -5.07 (95% CI, -5.99 to -4.15) for x-ray, and -5.0 (95% CI, -6.0 to -4.0) for MRI.
  • The pooled error rate for any error in LLM-rewritten reports was 7.2% (95% CI, 5.1%-10.0%), with clinically significant errors occurring at a rate of 0.9% (95% CI, 0.6%-1.5%), highlighting the need for human oversight before release to patients.

IN PRACTICE:

"By adopting a careful, evidence-based approach, LLM-rewritten reports could evolve from a technical novelty into a cornerstone of patient communication," the authors wrote.

SOURCE:

This study was led by Samer Alabed, PhD, School of Medicine and Population Health, Institute for In Silico Medicine, National Institute for Health and Care Research, University of Sheffield, Sheffield, England. It was published online on February 15, 2026, in The Lancet Digital Health.

LIMITATIONS:

Most studies were small single–centre studies and skewed towards younger, English‑speaking, more educated participants, limiting generalisability. Reproducibility and transparency were poor — datasets and code were not shared, and prompting methods varied — and heterogeneity across meta-analyses was very high. Outcomes relied on self‑reported rather than objective understanding; no studies involved patients in the design of the simplified reports or evaluated real‑world use. The studies focused mainly on CT and MRI, limiting the applicability to other modalities.

DISCLOSURES:

This study received support from the National Institute for Health and Care Research Sheffield Biomedical Research Centre. The authors declared having no relevant conflicts of interest.

SUGGESTED FOR YOU

This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.

References


Share This Article

Comments

Leave a comment