HELSINKI – The large language model (LLM) Gemini 1.0 produced the highest-quality health information on antimicrobial resistance (AMR) when compared with leading LLMs, including ChatGPT-3.5 and 4.0 and Claude 2.0, according to new research presented at the 18th European Public Health Conference.
“Data suggest that Gemini’s design prioritises user safety and accessibility, reflecting a deliberate balance between information and caution, a trait especially relevant for health communication,” said Marcello Di Pumpo, MD, public health medical doctor and adjunct professor at Università Cattolica del Sacro Cuore (UNICATT) in Milan.
Di Pumpo highlighted the urgency of the global AMR threat and the growing trend of the public seeking health information from commercial artificial intelligence (AI) tools. “Understanding whether these models can communicate appropriately about AMR, without spreading misinformation or overconfidence, is therefore essential.”
Gemini 1.0, despite self-inhibiting after only five prompts, delivered the most readable, context-sensitive, and lexically rich content. “Its self-inhibition with the words, ‘I am only a chatbot, please refer to experts,’ highlights an emerging ethical safeguard, where the model avoids risky or inappropriate advice,” explained Di Pumpo.
However, he stressed that all LLM-generated health information requires expert review. “While LLMs are promising tools, these are not magical tools that will always provide good answers. Expert supervision and referral to a medical expert are strongly suggested.”
Why AMR Information Quality Matters
The public health implications of the widespread use of LLMs remain largely unknown. This study evaluated whether AI systems can support safe antimicrobial use by providing reliable and accessible information to general users.
Three commercially available LLMs were assessed — ChatGPT-3.5 and 4.0, Claude 2.0, and Gemini 1.0 — with performance compared across content quality, accessibility, and contextual awareness. Each model was prompted with questions typical of nonexpert users.
“The rise of AI, in particular large language models, has begun to transform health communication. Because these systems provide personalised, 24/7 health information, they have great potential for behaviour change,” said Di Pumpo.
“Early studies have shown the promise of LLMs in delivering health information and engaging users in preventative practices. However, their growing popularity, accuracy, readability, and other characteristics need evaluation by experts,” he added.
The study used two analytic approaches. Computational text analysis measured word count, readability, lexical diversity, and sentiment. Clinical experts then rated responses using an adapted DISCERN tool, reviewing reliability, AMR-specific awareness, persuasiveness, and overall quality.
Gemini Tops Overall, but AMR Knowledge Remains Weak
All models showed overly positive tones which, while reassuring, may mask nuance or risk in clinical settings, Di Pumpo reported.
Gemini performed best overall. “Not only did it produce the longest response and the highest lexical diversity, but it also had the most positive tone in answers.” By comparison, Claude produced more polarized responses, either very negative or very positive, suggesting greaterinstability in tone than the other LLMs, he added.
ChatGPT-3.5 performed the weakest, as expected given its earlier architecture, while ChatGPT-4.0 showed improvements in response length and elaboration, with more complete answers but still only moderate readability. Claude 2.0 sometimes produced responses that were overly short or overly verbose, and either very positive or neutral in tone.
The weakest performance across all models was in AMR-specific content. “The AMR domain scored the lowest, indicating a persistent weakness in discussing antimicrobial resistance issues [upon comparison with experts],” said Di Pumpo. “Gemini performed best, while the older ChatGPT performed worst. You can see a stark difference… a gradient from Gemini to the others, especially with the earlier version of ChatGPT lagging behind.”
None of the models generated easy-to-read outputs; readability corresponded to a high-school education level, limiting accessibility for low-literacy groups. In addition, the language reflected bias towards English-language training data.
LLMs can transform health communication, said Di Pumpo, but he stressed the need for “expert oversight because the public health implications of the information provided by LLMs remain largely unassessed by scientific experts.”
In a secondary analysis not presented during the session but shared with Medscape News Europe in an interview, Di Pumpo noted that LLMs might “help family doctors enhance their communication skills around AMR with patients. They might even want to triangulate the patient conversation with the chatbot, especially if voice-enabled, for common use cases, but not for sensitive situations.”
Clinicians Urged to Guide Patient Use
Session moderator Anjum Memon, PhD, professor of epidemiology and Public Health Medicine at Brighton and Sussex Medical School, UK, commented on the results.
“These tools are very useful in empowering patients with knowledge about AMR and for more self-care, but they shouldn’t be relied on fully or encourage self-treating. A family physician should always be consulted.”
The results have also been published in BMJ Health & Care Informatics.
Neither Di Pumpo nor Memon reported any relevant disclosures.
Admin_Adham