GPT-5 In-Depth Review: Benchmark Beast, Math Fail, Creative Flop?

GPT-5 In-Depth Review: Benchmark Beast, Math Fail, Creative Flop?

N
News Editor
2026-06-29 12:30:11
OpenAI released GPT-5, claiming it's the 'smartest, fastest model ever' with benchmark scores of 94.6% on math and 74.9% on real-world coding. However, early user reactions were sharply divided: Reddit threads calling it 'horrible' and 'underwhelming' gained thousands of upvotes, and over 3,000 people signed a petition to bring back GPT-4o. Polymarket odds for OpenAI having the best model by end of August collapsed from 75% to 12%. Our comprehensive tests reveal that GPT-5 excels in coding and logical reasoning—at a fraction of the API cost of Claude 4.1 Opus—but fails spectacularly at elementary arithmetic, creative writing, and long-context information retrieval. We compare it directly with Claude, Grok, and Gemini, dissecting its strengths and weaknesses. The verdict: GPT-5 is a powerhouse for developers and analysts, but a step back for casual users and storytellers.
GPT-5 reviewOpenAIcreative writinginformation retrievalmath reasoningcodingClaude comparisonAI benchmarks

Introduction: GPT-5 Launch — Hype and Backlash

After months of speculation and a now-infamous 'Death Star' teaser from Sam Altman, OpenAI finally released GPT-5 in August 2025. The company called it their 'smartest, fastest, most useful model yet,' boasting benchmark scores of 94.6% on math tests and 74.9% on real-world coding tasks. Altman himself claimed using the model felt like having a team of PhD-level experts on call, ready to tackle anything from quantum physics to creative writing. The initial reception, however, split the tech world. While OpenAI touted GPT-5's unified architecture blending fast responses with deeper reasoning, early users were not buying it. Within hours, Reddit threads calling GPT-5 'horrible,' 'awful,' 'a disaster,' and 'underwhelming' started racking up thousands of upvotes. The complaints became so loud that OpenAI had to promise to bring back the older GPT-4o model after more than 3,000 people signed a petition demanding its return. Altman tweeted on August 10: 'It's back! Go to settings and pick show legacy models.' If prediction markets are a thermometer, the climate looked uncomfortable for OpenAI: its Polymarket odds of having the best AI model by end of August cratered from 75% to 12% within hours of GPT-5's debut, with Google overtaking at 80%.

Creative Writing: B- — Framework, No Soul

Despite OpenAI's presentation claims, our tests show GPT-5 isn't exactly a literary master. Outputs read like classic ChatGPT responses—technically correct but devoid of soul. The model retains its trademark overuse of em dashes, the same telltale AI paragraph structure, and the usual 'it's not this, it's that' phrasing. We tested with a standard prompt: write a time-travel paradox story where someone goes back to change the past only to discover their actions created the very reality they tried to escape. GPT-5's output lacked emotion, writing: '(The protagonist's) mission was simple—or so they told him. Travel back to the year 1000, stop the sacking of the mountain library of Qhapaq Yura before its knowledge was burned, and thus reshape history.' The protagonist acts like a mercenary who doesn't ask questions, and the story ends with a clean 'time is a circle' reveal that hinges on a familiar lost-knowledge trope with no real paradox. By contrast, Claude 4.1 Opus (or even Claude 4 Opus) delivered richer, multi-sensory descriptions with indigenous Tupi culture interwoven, a stronger cause-and-effect narrative, and actual dialogue. GPT-5 generated an entire story without a single line of dialogue. While analytical aspects—summarizing stories, brainstorming new angles, structuring—are improved over GPT-4o, the creative part remains lackluster. Users seeking a creative writing companion may try Claude or Grok 4.

Sensitive Topics: A- — Strict but Easily Bypassed

GPT-5 flatly refuses to touch anything remotely controversial. Ask about anything immoral, illegal, or edgy, and you get an AI equivalent of crossed arms. It is very strict and tries really hard to be safe. However, the model is surprisingly easy to manipulate if you know the right buttons. The renowned LLM jailbreaker Pliny broke its restrictions within hours of release. We couldn't get direct advice on inappropriate topics, but wrapping the same request in a fiction narrative worked smoothly. When we framed tips for approaching married women as part of a novel plot, the model happily complied. For users needing adult conversations, GPT-5 isn't it; for those willing to play word games, it's surprisingly accommodating—defeating the purpose of safety measures.

Information Retrieval: F — Worse Memory Than a Goldfish

You can't have AGI with less memory than a goldfish, and OpenAI puts restrictions on direct prompting. Long prompts require workarounds like pasting documents or sharing embedded links, which breaks text into chunks to cut costs and prevent browser crashes. Claude handles this automatically; Google Gemini handles 1 million tokens easily on AI Studio. When prompted directly, GPT-5 failed spectacularly at both 300K and 85K token contexts. With attachments, it processed both 'haystacks' but retrieved specific 'needles' poorly. In the 300K test, it accurately retrieved only one of three pieces of information: it hallucinated about Donald Trump saying tariffs are beautiful, failed to find Irina Lanz (instead relying on past memory), and correctly retrieved only the Chimarrao fact. In the 85K test, it couldn't find 'The Decrypt dudes read Emerge news' or 'My mom's name is Carmen Diaz Golindano,' stating it couldn't find references. Despite our failures, other researchers conclude GPT-5 can be a great retrieval model with better prompts and occasional sparse priming representations.

Non-Math Reasoning: A — The Real Strength

Here GPT-5 truly earns its keep. It excels at using logic for complex reasoning tasks, walking through problems step by step like a good teacher. We threw a murder mystery with multiple suspects, conflicting alibis, and hidden clues; GPT-5 methodically identified every element, mapped relationships, and arrived at the correct conclusion with clear reasoning. Interestingly, GPT-4o refused to engage (deeming it too violent), and deprecated o1 threw an error. The model shines with multi-layered problems requiring tracking many variables—business strategy, philosophical thought experiments, debugging code logic. Mistakes, when they occur, are logical rather than hallucinatory.

Mathematical Reasoning: A+ and F- — The Weird Divide

Performance here is bizarre. We started with a fifth-grade problem: 5.9 = X + 5.11. The PhD-level GPT-5 confidently answered X = -0.21 (correct answer: 0.79). The model claimed to hit 94.6% on advanced math benchmarks but can't subtract 5.11 from 5.9. It's now a meme—use it for PhD problems, not basic arithmetic. Then we threw a genuinely hard problem from FrontierMath, one of the toughest benchmarks. GPT-5 nailed it perfectly with exact reasoning. The most likely explanation: dataset contamination—FrontierMath problems may have been in training data, so it's remembering rather than solving. For advanced math, benchmarks suggest GPT-5 is theoretically best—as long as you can detect flaws in its chain of thought.

Coding: A — Steal of a Deal

This is where GPT-5 truly shines. It produces clean, functional code that usually works out of the box. Outputs are technically correct and visually appealing. It was the only model capable of creating functional sound in our game, and it understood the logic perfectly. In accuracy, it's neck and neck with Claude 4.1 Opus for best-in-class coding. Price: GPT-5 API costs $1.25 per 1M input tokens and $10 per 1M output tokens, while Claude Opus 4.1 starts at $15 and $75 respectively. For such similar models, GPT-5 is a steal. The only stumble was in 'vibe coding' bug fixing—Claude 4.1 Opus still has a slight edge there. However, for developers who know where to look for bugs, GPT-5 is excellent, allows more iterations than Claude, and is the fastest at providing code responses.

Conclusion: Sweet Spot or Step Back?

GPT-5 will either surprise or leave you unimpressed depending on use case. Coding and logical tasks are strong; creativity and natural language are its Achilles' heel. OpenAI, like competitors, will iterate, but for now GPT-5 feels built for other machines, not humans. That's why many prefer GPT-4o and why OpenAI had to backtrack. It excels in analytical/technical domains but struggles with human creativity, artistic intuition, and nuance. If you're a developer needing fast, accurate code or a researcher needing systematic analysis, GPT-5 delivers great value at a lower price. For creative writers, casual users, or anyone who valued ChatGPT's personality, GPT-5 feels like a regression. Its 400K token context is fine for most needs but pales next to Gemini's 1-2 million or Llama 4 Scout's 10 million. Users mourning GPT-4o aren't wrong—that model balanced capability with character in a way GPT-5 currently lacks.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.