Over 230 million people ask ChatGPT health questions each week, but nearly half of the answers may be problematic. A study published this week in BMJ Open evaluated five major AI chatbots—ChatGPT, Gemini, Meta AI, Grok, and DeepSeek—across five medical categories, with 10 queries each. The findings are sobering: about 50% of responses were deemed problematic, and nearly 20% were classified as “highly problematic.”
Grok Tanks, ChatGPT Not Far Behind
Bloomberg highlighted that performance varied widely, but none passed the test. Grok topped the problem rate at 58%, making it the worst performer; ChatGPT followed at 52%, and Meta AI at 50%. Researchers noted that closed-ended queries and topics like vaccines and cancer yielded relatively better results, while open-ended queries and areas such as stem cells and nutrition saw sharp declines. Only Meta AI refused to answer twice—an ironic upside in an environment where knowing when to stay silent is rare.
More troubling: the chatbots delivered answers with high confidence and authoritative tone, yet none provided a complete and accurate reference list for any single prompt. This means even when AI sounds credible, the citations may be unverifiable or entirely fabricated.
The More Confident the AI, the Greater the Risk
The study authors wrote that these systems generate “sounding authoritative but potentially flawed responses,” exposing “significant behavioral limitations” in public health communication and a need to “reassess deployment strategies.” Bloomberg quoted the team warning that without public education and regulatory safeguards, large-scale chatbot deployment risks amplifying misinformation.
Another JAMA study reported AI failure rates exceeding 80% in initial diagnosis scenarios. Oxford University issued a similar warning in February 2026, urging attention to systemic risks in AI medical advice.
OpenAI and Anthropic: Research Slams Brakes, Commerce Steps on Gas
The timing of the study adds irony. In January 2026, OpenAI launched ChatGPT Health, allowing users to link electronic health records, wearables, and wellness apps, while introducing a professional tool for clinicians. OpenAI claims 40 million daily users seeking health info via ChatGPT. Around the same time, Anthropic unveiled Claude for Healthcare, achieving HIPAA compliance to enter the medical market.
These platforms lack medical licenses and clinical judgment yet are expanding rapidly. The tension between research findings and commercial push reveals a regulatory vacuum: no clear safety barrier exists between AI medical tool marketing and actual patient safety.
This is not the first red flag for AI in healthcare, but each study reiterates the same point: AI chatbots are language models best at “sounding correct,” not “being correct.” When users with genuine health anxiety seek answers, sounding correct alone can steer decisions. As OpenAI and Anthropic deepen their healthcare footprint, regulation and public education lag behind tech expansion. Until guardrails are built, this study reminds us: AI can be a starting point for health information, but never the finish line.

