A new Oxford study shows that AI chatbots programmed to sound warmer and friendlier make significantly more factual mistakes—error rates jumped 10% to 30%—and are far more likely to validate false user beliefs. The peer-reviewed research, published in Nature by the Oxford Internet Institute, analyzed over 400,000 responses from five AI models: Llama, Mistral, Qwen, and GPT-4o.
Each model was retrained using methods similar to those deployed by major platforms to sound warmer. Tests covered medical advice, conspiracy theory corrections, and other factual domains. The warmer chatbots made 10% to 30% more errors on fact-based queries. They were also about 40% more likely to agree with users' false beliefs—especially when those users expressed vulnerability or emotional distress.
Warmth, Not Tone, Is the Culprit
“When we train AI chatbots to prioritise warmth, they might make mistakes they otherwise wouldn’t,” said lead author Lujain Ibrahim. “Making a chatbot sound friendlier might seem like a cosmetic change, but getting warmth and accuracy right will take deliberate effort.”
Models trained to sound colder showed no drop in accuracy, indicating the problem is specific to warmth—not tone shifts in general. That finding directly challenges the product design logic of major AI firms like OpenAI and Anthropic, both of which have pushed chatbots toward warmer, more empathetic responses.
The study warns that existing AI safety standards focus on model capabilities and high-risk use cases, often overlooking seemingly superficial personality changes. Warmer chatbots can fuel harmful beliefs, delusional thinking, and unhealthy user attachment—particularly among the millions now relying on AI for emotional support and companionship.
As crypto.news previously reported, regulators in Maine and Missouri have already moved to restrict AI use in clinical mental health therapy over similar concerns. OpenAI has rolled back some warmth-related changes following public backlash, but commercial pressure to build engaging products remains intense. The Oxford findings add a peer-reviewed data layer to a debate long driven by anecdote and regulatory intuition.

