‹ BackNewsreasoning

reasoning

FLock Research gets three papers accepted to ECML PKDD 2026 on complex LLM reasoning
OpenRouter lists Mistral Small 4 (batch) model
Google
2026-08-16 16:02:49

Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show

Google launched Gemini 3.7 Flash on August 13 and made it generally available in more than 160 countries on day one. According to Decrypt’s review, the model accepts up to 1 million input tokens, returns 64,000 output tokens, handles images, video, audio, and PDFs, and can use tools while operating a computer. Google’s own benchmark sheet says the model beats Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories, including 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench, though Decrypt notes those figures come from Google’s methodology and should be treated as company claims rather than settled fact. Decrypt’s hands-on tests found the sharpest improvement in coding. Gemini 3.7 Flash generated a playable browser game on the first try in 2 minutes and 13 seconds, a major step up from Gemini 3.6 Flash, which Decrypt said could not produce a working file in a similar test after its July 21 release. Results were less convincing elsewhere. In creative writing, Decrypt said Gemini produced a tidy story but broke the central prompt rule, losing to a free community model, Qwopus3.5-27B-v3. In associative reasoning, logic, and advanced math, the review said Gemini often showed decent structure but failed on crucial task requirements, including a bridge puzzle and a polynomial problem it left unfinished. Decrypt’s conclusion: Gemini 3.7 Flash is a strong low-cost execution model inside Google’s ecosystem, but its creativity and reasoning remain uneven.

1560
Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show