BackLlama 4

Llama 4

GPT-5 review
2026-06-29 12:30:11

GPT-5 In-Depth Review: Benchmark Beast, Math Fail, Creative Flop?

OpenAI released GPT-5, claiming it's the 'smartest, fastest model ever' with benchmark scores of 94.6% on math and 74.9% on real-world coding. However, early user reactions were sharply divided: Reddit threads calling it 'horrible' and 'underwhelming' gained thousands of upvotes, and over 3,000 people signed a petition to bring back GPT-4o. Polymarket odds for OpenAI having the best model by end of August collapsed from 75% to 12%. Our comprehensive tests reveal that GPT-5 excels in coding and logical reasoning—at a fraction of the API cost of Claude 4.1 Opus—but fails spectacularly at elementary arithmetic, creative writing, and long-context information retrieval. We compare it directly with Claude, Grok, and Gemini, dissecting its strengths and weaknesses. The verdict: GPT-5 is a powerhouse for developers and analysts, but a step back for casual users and storytellers.

190
GPT-5 In-Depth Review: Benchmark Beast, Math Fail, Creative Flop?