The recent launch of ChatGPT Health by OpenAI, which made headlines in recent weeks, followed shortly by Claude for Healthcare from Anthropic, raises critical questions for clinicians about reliability, transparency, privacy, and compliance with European law.
ChatGPT Health and Claude for Healthcare were presented as new health management experiences designed to securely integrate personal health information with generative AI. The aim is to help users feel better informed, more prepared, and more aware of how to manage their health.
The initiative builds on the recognition that health is already one of the most common areas of use for ChatGPT, with an estimated hundreds of millions of users asking questions about health and wellness each week.
According to OpenAI, users can connect electronic health records, health applications, wearable devices, and medical documents to the system. The platform then generates personalized, contextualized responses intended to help users interpret laboratory results, prepare for medical appointments, receive guidance on nutrition and physical activity, and evaluate insurance options, which is an issue of relevance in the US.
Scientific Validation
OpenAI states on its website that ChatGPT Health adheres to the same high standards of privacy and security as ChatGPT, reinforced by additional measures for health data, including dedicated encryption systems and isolation mechanisms designed to protect the confidentiality of the conversations.
However, no peer-reviewed studies have demonstrated the clinical robustness of the system. OpenAI reported that approximately 260 physicians across 60 countries tested the system and provided feedback on hundreds of thousands of responses. Nevertheless, no published scientific report has detailed the methodology, outcome measures, or training data in accordance with accepted clinical research standards.
Second, ChatGPT Health is presented as a purely informational tool; however, it functions like a medical device. When people turn to these systems for information, as they have long done with Dr Google, health apps, and other digital platforms, they are often looking for guidance to help them make decisions regarding diagnosis or treatment. In this context, what is described as an information tool effectively becomes a decision support tool for citizens and patients.
A recent survey indicated that 63.9% of Italians attending a medical visit used online information to verify the diagnosis or therapy recommended by their physician. In two out of three cases, respondents reported questioning the recommendations they received.
Studies evaluating general-purpose generative artificial intelligence (AI) systems for health-related queries consistently show that responses may appear credible but contain inaccuracies upon closer review. In some instances, these errors reflect the limitations of the scientific reliability of the sources used for training. Google has recently restricted certain outputs from its AI Overview system when responding to medical-related questions.
Privacy and Regulation
Therefore, privacy remains a major concern. OpenAI states that user data, including conversations, uploaded medical reports, electronic health record content, and data from health applications, are not used for training or secondary purposes. Nonetheless, concerns persist regarding potential misuse, particularly if the systems are not fully compliant with the EU General Data Protection Regulation.
The issue does not lie solely in noncompliance with the General Data Protection Regulation. A more substantial concern is the failure to comply with the principles set out in the EU Artificial Intelligence Act. The EU Artificial Intelligence Act requires that AI systems operating in Europe ensure transparency and explainability and that their outputs are replicable.
In short, ChatGPT Health is not authorized in any country to operate as a medical device, and it would face difficulties operating in Europe, even as a general AI system.
Old Questions, New Tools
On closer examination, the situation is not very different from the era of “Dr Google.” The way Google and ChatGPT Health respond to medical questions and the sources they draw on have not fundamentally changed. What has changed is the tone. Generative AI tends to use fluent and confident language that can sound reassuring and authoritative, which may make users more inclined to accept the answer as plausible.
Simultaneously, it is important to recognize the limitations of these systems. AI tools do not accurately interpret clinical data or reasons by applying logical rules. They generated responses by predicting which words or phrases were statistically most likely to follow prompts. ChatGPT Health operates on the same underlying principle.
Here is a revised version with a more measured, neutral tone and without sensational wording.
Even digital health solutions that lack solid scientific validation often fail to produce meaningful results. Other initiatives launched by major technology companies began with ambitious objectives but were later scaled back or gradually phased out. Examples include Google Health, Microsoft HealthVault, and Medpedia, none of which achieved a sustained impact.
Using Generative AI in Healthcare: Not All Efforts Are in Vain
This does not imply that generative AI cannot contribute meaningfully to healthcare. However, its value depends on the way it is developed and validated. Several systems are already in use, including MedGemma and Med-PaLM, developed by Google, which are trained on peer-reviewed scientific literature indexed in bibliographic databases such as Medline and are intended to support clinical decision-making.
Similarly, the Open Evidence project draws on research published in journals such as the JAMA and The New England Journal of Medicine to assist physicians in making evidence-based decisions. For the public, the World Health Organization developed Sarah, a virtual coach trained on materials produced by the World Health Organization to promote healthier lifestyles and prevent noncommunicable diseases.
These initiatives are generally part of structured public health programs overseen by clinicians and researchers and are evaluated through studies using established research methodologies.
In the 1990s, Alessandro Liberati, one of the leading advocates of evidence-based medicine in Italy, asked, “Where is the evidence?” That question remains relevant today. It may be worth asking more frequently, especially when digital tools that claim to be intelligent propose solutions that can influence public health.
ChatGPT Health is not currently available in Italy, although users may register for a waiting list for Eugenio Santoro pending future availability.
Eugenio Santoro is a digital health researcher at the Mario Negri Institute for Pharmacological Research IRCCS in Milan, Italy, where he works in the Laboratory of Medical Informatics and in research activities focused on digital health and digital therapeutics. He is a member of the ICT group of the Federazione Nazionale degli Ordini dei Medici Chirurghi e degli Odontoiatri, a member of the Technical Scientific Committee of the Parliamentary Intergroup on Digital Health and Digital Therapeutics, the Emilia-Romagna Region’s “Gruppo di Lavoro IA in sanità” (AI in Healthcare Working Group), and a member of the Control Committee of the Istituto di Autodisciplina Pubblicitaria. He also serves on the Medical Device Observatory Software as a Medical Device at the Istituto Superiore di Sanità, Rome, Italy.
He has authored several books and numerous scientific articles in the field of digital healthcare. In 2021, he edited the Digital Health entry in Appendix X of the Treccani Encyclopedia, dedicated to 21st century terminology. He teaches in several university master’s degree programs.
This story was translated from Univadis Italy, part of the Medscape Professional Network.
Admin_Adham