DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race

N
News Editor
2026-08-12 23:48:41
DeepSeek V4 Pro and xAI’s Grok 4.6 arrived on the same day, setting up a direct comparison across benchmark scores, pricing, and real task execution. DeepSeek V4 Pro posted strong results in CyberGym, AutomationBench, Terminal-Bench 2.1, and DeepSWE, in several cases beating Opus 4.8 and narrowing the gap with Fable 5. Grok 4.6, meanwhile, kept pace in broader capability rankings and moved ahead in several coding and knowledge-work evaluations, including GDPval-AA v2, AA-Briefcase, and Harvey LAB. Pricing became another focal point. DeepSeek listed output pricing at $0.87 per million tokens, compared with $6 for Grok 4.6, $30 for GPT-5.6 Sol, $25 for Claude Opus 5, and $50 for Fable 5. Early hands-on tests cited in the source showed the two models trading wins across website creation, design generation, frontend work, and game-building tasks. In one Flappy Bird comparison, DeepSeek used more than 20,000 tokens at a cost of $0.019, while Grok used about 5,000 tokens at a cost of $0.03. The resulting DeepSeek version was described as more polished. The source’s overall takeaway was that the two newly released models are now operating at roughly the same level.

DeepSeek V4 Pro and Elon Musk-backed Grok 4.6 were released on the same day, turning model comparisons into a head-to-head contest centered on one shared goal: handling long tasks by calling tools, editing code, checking outputs, and delivering something usable at the end.

Benchmark scores put both models in direct competition

DeepSeek V4 Pro’s headline numbers came from a string of agent-focused evaluations.

In the CyberGym cybersecurity agent test, DeepSeek V4 Pro scored 83.3, ahead of Fable 5 at 83.1 and Opus 4.8 at 78.3. In AutomationBench, it posted 31.8, topping Fable 5 at 29.1 and Opus 4.8 at 27.2.

On Terminal-Bench 2.1, DeepSeek V4 Pro reached 87.9. That was above Opus 4.8 at 85.0 and only 0.1 points behind Fable 5 at 88.0.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 3

Other hard agent tests also showed DeepSeek V4 Pro moving past Opus 4.8. In the tool-enabled “Humanity’s Last Exam,” it scored 60.0 versus 57.9 for Opus 4.8, while Fable 5 stood at 63.0. Its software-engineering agent result on DeepSWE jumped from 12.8 in the preview version to 62.7, nearly 4.9 times higher. That put it above Opus 4.8 at 58.0 and 7.3 points behind Fable 5 at 70.0.

Grok 4.6 was positioned against Fable 5 and GPT-5.6 Sol.

On a composite intelligence index, Grok 4.6 scored 61, one point behind Fable 5 at 62. In coding tests, it reached 69.9% on CursorBench, ahead of GPT-5.6 at 67.2% and 0.6 percentage points behind Fable 5 at 70.5%. On FrontierCode, Grok 4.6 scored 61.3%, above GPT-5.6 at 60.6% and below Fable 5 at 63.6%.

In knowledge-work evaluations that the source described as closer to real workplace delivery, Grok 4.6 ranked first across all three tests cited. It scored 1753 Elo on GDPval-AA v2, beating Fable 5 at 1741 and GPT-5.6 at 1728. On AA-Briefcase, it reached 1577 Elo, ahead of Fable 5 at 1574 and GPT-5.6 at 1502. On the legal task benchmark Harvey LAB, Grok 4.6 scored 15.8%, compared with 11.3% for Fable 5 and 2.5% for GPT-5.6.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 4

Pricing moved sharply lower

Cost was a major part of the release-day comparison.

Per 1 million output tokens, Grok 4.6 was priced at $6, GPT-5.6 Sol at $30, Claude Opus 5 at $25, and Fable 5 at $50. DeepSeek V4 Pro was listed at $0.87. The source said that puts DeepSeek at about one-seventh the cost of Grok 4.6, one-thirty-fifth the cost of GPT-5.6 Sol, one-twenty-ninth the cost of Claude Opus 5, and one-fifty-seventh the cost of Fable 5.

Early hands-on tests moved from benchmarks to real tasks

The first wave of real-world testing has already started, with developers pushing both models beyond scoreboards.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 5

In one initial run, DeepSeek V4 Pro built a complete interactive 3D Earth in a browser from a single prompt. Basic interactions such as drag, rotation, and zoom were available. Global data flows, animated routes, and geographic markers were laid across the globe. The source also pointed to atmospheric scattering, cloud rendering, day-night lighting, and a full user interface.

AI blogger “Xiangyang Qiaomu” then ran three small tasks with DeepSeek V4 Pro and used the same prompts for two of them with Grok 4.6 to compare the final outputs.

The first task involved calling three Skills to develop and deploy a website. DeepSeek V4 Pro completed the full process, and the source described its page design and overall finish as steady. The second task was to recreate 60 design styles and present them in a full set of Bento cards. DeepSeek differentiated fonts, color schemes, and layouts clearly, with a strong overall visual result. The last task was to build a 3D brick-breaker game from scratch. The finished game was playable and included a 3D scene, background music, and dynamic sound effects.

When the same prompts were given to Grok 4.6, the 3D brick-breaker task ended up close to even. According to the source, Grok’s game matched DeepSeek V4 Pro on visual quality, scene completeness, and playability. On the 60-style Bento task, Grok 4.6 recovered ground, with some pages showing more mature layout, color use, and visual hierarchy.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 6

Flappy Bird and frontend tests produced split outcomes

A sharper comparison came from a Flappy Bird task. Developer Jun Song used the same prompt and asked DeepSeek V4 Pro and Grok 4.6 to build the game from scratch.

The two models took very different paths. DeepSeek V4 Pro used more than 20,000 tokens at a cost of $0.019. Grok 4.6 used about 5,000 tokens, but the cost came to $0.03.

On the final result, the source said DeepSeek V4 Pro had the edge. Its game included distant mountain scenery and layered clouds. The pipes had gradients and visible volume, and a floating “+1” animation appeared when the character passed through them. The source described the scene layering, control feedback, and end-to-end completion as stronger than Grok 4.6.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 7

In another frontend test conducted by developer Hamza, DeepSeek V4 Pro again finished ahead of Grok 4.6 in both page completeness and visual output.

That did not mean DeepSeek V4 Pro won every comparison. In one pelican test set mentioned by the source, V4 Pro delivered a more complete overall image than the Flash version, and the elephant design looked better, but the pelican’s movement direction was completely reversed.

The gap became clearer in another test that asked DeepSeek V4 Pro, GPT-5.6 Sol, and Claude Opus 5 to generate the same cherry blossom tree with Three.js. The source said V4 Pro trailed the other two top-tier models in trunk and foliage detail, lighting, depth of field, and overall atmosphere.

Same-day releases tightened the race with top AI labs

After several rounds of benchmark results and hands-on tests, the source’s conclusion was that DeepSeek V4 Pro and Grok 4.6 are now operating on roughly the same level.

DeepSeek V4 Pro and Grok 4.6 launch on the same day as early tests show a tight race 8

Rick De Oliveira described the two-model breakout with a short verdict: 「强大,而且实惠。」

The source argued that from cybersecurity and knowledge work to agent-led problem solving, both models are pushing toward frontier-level capability at lower cost. It also said Musk has already previewed Grok 4.7 with the stated goal of surpassing all top AI systems, while OpenAI and Anthropic are expected to answer with their own next moves.

The references listed in the source include DeepSeek’s pricing page, xAI’s news page, and test posts on X from vista8 and Jun Song. The original Chinese article was credited to the WeChat public account Xinzhiyuan, written by ASI Qishilu and edited by Moses and Taozi.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
500

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.