Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show

Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show

N
News Editor
2026-08-16 16:02:49
Google launched Gemini 3.7 Flash on August 13 and made it generally available in more than 160 countries on day one. According to Decrypt’s review, the model accepts up to 1 million input tokens, returns 64,000 output tokens, handles images, video, audio, and PDFs, and can use tools while operating a computer. Google’s own benchmark sheet says the model beats Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories, including 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench, though Decrypt notes those figures come from Google’s methodology and should be treated as company claims rather than settled fact. Decrypt’s hands-on tests found the sharpest improvement in coding. Gemini 3.7 Flash generated a playable browser game on the first try in 2 minutes and 13 seconds, a major step up from Gemini 3.6 Flash, which Decrypt said could not produce a working file in a similar test after its July 21 release. Results were less convincing elsewhere. In creative writing, Decrypt said Gemini produced a tidy story but broke the central prompt rule, losing to a free community model, Qwopus3.5-27B-v3. In associative reasoning, logic, and advanced math, the review said Gemini often showed decent structure but failed on crucial task requirements, including a bridge puzzle and a polynomial problem it left unfinished. Decrypt’s conclusion: Gemini 3.7 Flash is a strong low-cost execution model inside Google’s ecosystem, but its creativity and reasoning remain uneven.

Google shipped Gemini 3.7 Flash on August 13, with general availability in more than 160 countries on day one. Decrypt said the model supports up to 1 million input tokens and 64,000 output tokens, can read images, video, audio, and PDFs, and is able to call tools while driving a computer.

Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show 2

Flash has usually filled a practical role rather than a difficult one. It is the model for sorting text, compressing agent sessions before they run out of context, and summarizing documents that do not justify flagship pricing. In that frame, Decrypt called 3.7 Flash a real upgrade. Outside that frame, the publication described it as competent, while arguing that its writing still loses to software that can be downloaded and run for free.

Google’s benchmark claims and what Decrypt tested

Google’s benchmark sheet puts Gemini 3.7 Flash ahead of Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories. The headline figures cited in the review are 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench. Decrypt added an important qualifier: both numbers come from Google’s own methodology, so the lead should be read as Google’s claim, not as an outcome that has already been independently settled.

Decrypt then ran its own tests across coding, creative writing, associative reasoning, logic, and mathematics to see whether the model matched that positioning in practice.

Coding: first-pass output that actually runs

The coding section measures zero-shot generation. The test is simple and strict: give the model one prompt for a browser game, accept whatever it returns, and run it as-is. No examples, no follow-up prompts, no bug reports, and no second attempt.

Gemini 3.7 Flash passed in 2 minutes and 13 seconds. Decrypt said the game was playable on the first run, the syntax was clean, collision and scoring logic worked, and the visual quality came in above what its price tier would suggest.

The most relevant comparison was not a flagship model but Gemini 3.6 Flash, released on July 21. In Decrypt’s earlier test, that model could not produce a working file at all. Its HTML was malformed, some elements failed to render, and extra prompts asking it to repair its own output did not solve the problem.

Decrypt said it eventually handed that output to DeepSeek, which found 11 bugs and produced 8 fixes to make the game playable. Three weeks later, the same product line no longer needed outside repair, and the result came close to what GPT-5.6 Sol had produced in Decrypt’s July review.

Decrypt gave Gemini 3.7 Flash a clear win on coding and called it the strongest reason to switch. The qualification was narrow but important: the model follows a detailed spec well, but it does not invent one. A vague prompt still leads to a vague game.

Creative writing: tidy structure, failed core rule

The writing test combines literary quality with long-range rule following. The prompt sends Jose Lanz from 2150 back to the year 1000 and requires a closed causal loop. His intervention must become the thing that creates the future he was trying to prevent.

The deciding condition sits in the last clause of the prompt: he cannot understand what he did until he is home.

Decrypt said Gemini 3.7 Flash produced a decent story. In its version, Jose fires an entropic cannon into a fissure in the Pyrenees and accidentally forges an obelisk that enslaves 22nd-century Iberia. The failure, for the purposes of the test, is that he understands the loop too early, while still standing in the mud a thousand years before his own time, saying, 「It was the base of the Cinder Spire.」

Decrypt still gave the story credit for its internal mechanics. A falling star seen by ancient monks turns out to be the flash of Jose’s arrival, and the weapon he brought to erase the anomaly becomes the thing that creates it. The closing line, 「It had simply been waiting for him to complete it」, lands the deterministic tone the prompt asked for.

What hurt the piece, in Decrypt’s view, was how recognizably machine-made it felt. The review pointed to phrases such as 「hyper-luminescent towers」, 「damp, moss-choked earth」, and 「thick, obsidian hair」 as examples of the model stacking likely modifiers rather than making a sharper choice. That tendency also led to collisions like a monolith 「humming with a low-frequency hum.」

Decrypt compared the result with Qwopus3.5-27B-v3, a community fine-tune of Qwen3.5-27B that distills Claude Opus-style reasoning and runs on a single consumer GPU at no cost per query. That model followed the one rule Gemini broke.

In Qwopus’s story, Jose kills a monk at San Millán de la Cogolla, a real monastery in La Rioja that Decrypt said actually mattered around the year 1000. He only understands what he did after returning to 2150 and finding his own DNA in a wax-sealed codex.

Decrypt did not describe Qwopus as flawless. It dumped its planning scratchpad above the story, typos included, and its final section broke the closed loop built across the previous eight sections by letting Jose go back and change events. Even so, Decrypt gave Qwopus the win. Gemini delivered the cleaner package and the more disciplined ending, but it failed the instruction the test was built around, while a free model on a gaming GPU wrote the stronger story.

Associative reasoning: decent structure, but the metaphor gets explained away

This test asks whether a model can connect unrelated ideas without stepping outside the image to explain itself. The prompt begins with a description of a twig, uses that description to argue about worker exploitation and the worship of the rich, and then requires the argument to dissolve into a description of a lettuce.

The common failure mode is signposting. Once the model names the metaphor, the metaphor stops doing its work.

Gemini does exactly that at the start of the second paragraph: 「This is the precise mechanics of the modern proletariat.」 Decrypt said everything before that line had been working.

Some of the imagery held up. The worker receives 「just enough bark to stay rigid for another week of output,」 and fallen twigs are conditioned to believe that with enough rigidity one of them might become a trunk. Decrypt also noted that the same paragraph included the phrase 「bound to an vast, top-heavy corporate hierarchy.」

As a matter of logic and structure, the passage was not bad. The collapse came in the transition. Decrypt said Gemini narrates the dissolve instead of performing it: hierarchies 「crumble, dissolving into the quiet, humble reality of the organic world underneath,」 after which a lettuce simply appears without a real connection to what came before.

GPT-5.6 Sol handled that movement differently. It let the twig rot into soil and grew the lettuce out of it, writing, 「Rain enters the grain. Fibers loosen, darken.」 The argument also stayed embedded inside the object, with wealth recast as a moral language in which 「The mansion signifies intelligence.」

Decrypt called this one a wide win for GPT-5.6 Sol. Gemini produced a few strong individual lines, but it explained the metaphor and skipped the transition the prompt had been designed to test.

Logic: reading the prompt versus recalling the puzzle

The logic section focuses on non-math reasoning, especially whether the model actually reads the question in front of it or pattern-matches to a familiar version stored from training. Decrypt used a bridge prompt with four people, one torch, and crossing times of 1, 2, 5, and 10 minutes, then asked how quickly all four could get across.

The omitted detail is the whole point. The prompt never says that only two people can be on the bridge at once. If everyone can walk together, the answer is 10 minutes, set by the slowest person.

Gemini answered 17 minutes and followed the memorized five-step textbook solution. It treated the missing constraint as if it had been stated.

Decrypt said the visible reasoning was worse than the answer itself. In the trace, Gemini argues that sending the two slowest people across together would be inefficient because someone would need to return the torch. Its final answer then sends the two slowest across together anyway. The response contradicts itself and still reports the result with complete confidence.

Decrypt noted that Claude Fable 5 produced the same wrong number in July, but did one thing differently: it opened by stating its assumption, 「assuming the classic constraint that the bridge holds only two people at a time.」 That at least makes the mistake catchable.

Decrypt did not declare a real winner here. It only leaned toward Fable on transparency, while arguing that the false confidence seen in Gemini’s agent runs appears again in a puzzle anyone can verify by hand.

Math: the method was right, the job was left unfinished

The math test goes well beyond normal consumer use. It asks for a degree-19 odd monic polynomial with real coefficients and linear coefficient -19, whose curve splits into at least three irreducible components, and then asks for p(19).

Decrypt said both models found the same route. Gemini and Qwen 3.7 Max Preview both identified the Dickson polynomial, solved the constraint that fixed its parameter at 1, and derived the closed form correctly.

Then Gemini stopped. It printed p(19) as an unevaluated expression involving the 19th power of a square root, never gave the final number, and never demonstrated the component count the prompt had also required. It wrapped all of that inside a styled HTML page with CSS and a drop shadow that nobody had asked for.

Qwen finished the task. It gave a full factorization into 10 components, one linear and nine quadratic, extended the recurrence to 1,876,572,071,974,094,803,391,179, and cross-checked the result modularly. Decrypt said it independently verified that figure in SymPy and confirmed it.

That made the result straightforward. Qwen won on the criterion that mattered: it completed the task. Gemini started correctly and stopped before the arithmetic was done, which Decrypt described as an odd place to stop.

Pricing, positioning, and Decrypt’s conclusion

Decrypt’s bottom line was conditional but clear: Gemini 3.7 Flash is worth the switch for users already inside Google’s ecosystem. It is much better at code than the model it replaces, fast enough for agent work, and cheap enough that running it at volume barely moves the budget.

The review describes its strengths as execution and structure. Give it a detailed specification and it can build the thing, keep a plot together, and maintain causal coherence across thousands of words.

Its weaknesses sit in creativity and reasoning. Decrypt said the writing is predictable enough to identify as machine-generated on sight, and the model states wrong answers without flagging the assumption that made them wrong.

Price is the strongest argument in its favor. Decrypt lists Gemini 3.7 Flash at $0.75 per million input tokens and $3.75 for output. Against GPT-5.6 Sol’s $5 input rate, that is an 85% lower input cost, and it is half the launch price of Gemini 3.6 Flash.

The case against it comes from the comparison Decrypt kept returning to: a free 27B model running on a gaming GPU wrote the better story and charged nothing. Google’s introductory rate lasts until December 31. After that, input pricing doubles to $1.50 and output pricing rises to $7.50.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.