user Admin_Adham
12th May, 2026 12:00 AM
Test

Language Models as Emergency Room Decision Support

A study recently published in Science suggests that large language models (LLMs) are at least on par with physicians when it comes to making diagnoses. The US research team, led by Peter Brodeur, MD, a resident physician at Beth Israel Deaconess Medical Center in Boston, reports that the difference between AI and medical professionals is largest in time-critical scenarios where clinicians must make decisions with limited patient information.

Brodeur and his team examined the capabilities of the language models OpenAI o1, OpenAI o1-preview, and GPT-4o. For the first part of the study, they used case catalogs and datasets intended, among other things, for education and training: These included diagnostic reports from NEJM as well as case studies derived from a podcast by the American College of Physicians. They also incorporated case vignettes from a “landmark” study evaluating computer-aided diagnostic systems from 1994 and a study from JAMA 2021.

The LLMs were asked to describe clinical cases, generate and prioritize differential diagnoses, recommend and interpret diagnostic tests, and justify their decisions. Their responses were compared with the original treatment plans and independently evaluated by two physicians; when the two disagreed, they discussed the case or a third reviewer made the final determination.

Emergency Triage Performance

In the second part of the study, researchers gave the LLMs and two physicians real, unstructured data drawn from a US emergency department. Available information varied by case — sometimes symptoms and vital signs, other times medical history. Some cases involved new arrivals to the emergency department, whereas others were existing patients. Using the information at hand, the clinicians and the model determined the appropriate course of action across points of care, from the initial emergency department visit through admission to the ICU. Two internists who were co-authors of the study then evaluated the responses in a blinded review.

Although the tested language models did not always provide completely correct diagnoses, their outputs were rated as at least helpful in most cases. Most of the suggested tests were also deemed appropriate. In the current study, the tested models also outperformed medical professionals in diagnostic performance and decision-making at the various points of contact in the emergency department. The advantage of AI was particularly pronounced at the first point of contact, namely admission to the emergency department.

SUGGESTED FOR YOU

“Overall, the study is methodologically stronger than many earlier benchmark studies on AI. It combines different clinical tasks with a human baseline and partially blinded evaluations, as well as real emergency department cases. This significantly increases the study’s validity. It convincingly demonstrates that modern language models can now achieve a very high level of performance in text-based medical diagnostic tasks,” commented Gitta Kutyniok, chair of Mathematical Foundations of Artificial Intelligence, Ludwig-Maximilians-Universität München, München, Germany, on the results to the Science Media Center.

Text-Only Study Limits

Felix Nensa, research group leader at the Institute for Artificial Intelligence in Medicine and senior attending physician at the Institute of Diagnostic and Interventional Radiology and Neuroradiology at Essen University Hospital in Essen, Germany, also confirms that the study is methodologically sound. Nevertheless, according to Nensa, the practical significance is limited as the tested models — OpenAI o1-preview and OpenAI o1 — were already replaced by o3 in early 2025. Furthermore, for clinical tasks, agent-based systems are increasingly being used that access specialized knowledge databases and are generally far superior to pure language models. “Limiting the study to text-based queries also falls short. Modern models operate multimodally; they can process images and videos in particular. This makes a significant difference when it comes to medical questions,” emphasized Nensa. The practical utility of such benchmarks therefore remains limited.

Thomas Neumuth, technical director at the Innovation Center for Computer Assisted Surgery at the University of Leipzig in Leipzig, Germany, also sees weaknesses: “Some sub-experiments use only five or six cases, and the ‘right or wrong’ evaluation depends on medical judgment. Furthermore, only text was tested, not what actually happens in everyday clinical practice.” Neumuth therefore considers the results to be only partially applicable. “In everyday clinical practice, much more happens than just reading text: Doctors observe whether a patient seems restless, listen to their breathing, look at x-rays, and ask follow-up questions. All of that is completely missing from the study because the model is only given fully documented cases.” Neumuth also believes that the task of providing a second opinion at three fixed points does not accurately reflect a real emergency room. After all, the main focus is on rapid triage.

Clinician AI Collaboration

Kutyniok stressed that pairing physicians with AI is currently seen as the most promising path — but only when that collaboration is thoughtfully designed. Partnership alone won’t automatically improve outcomes: If clinicians accept AI recommendations uncritically, automation bias can take hold. “Only if they use the AI as a structured second opinion can quality ultimately improve,” Kutyniok added.

Whether an AI system truly helps cannot be determined by endless new desk-based tests, but only through real clinical trials, Neumuth emphasized. These should measure what matters: fewer misdiagnoses, shorter wait times, and better patient outcomes. And ongoing monitoring during use is needed, similar to what is done with newly approved medications.

This story was translated from Medscape’s German edition.


Share This Article

Comments

Leave a comment