DeepSeek’s V4.1 Flash nearly matches GPT-6 Astra in design benchmark at roughly 1.4% of the cost

DeepSeek’s V4.1 Flash nearly matches GPT-6 Astra in design benchmark at roughly 1.4% of the cost

N
News Editor
2026-09-10 21:16:03
OpenDesign, the company behind OpenDesign Arena, tested 13 AI models on the same batch of practical design tasks this week, focusing on work such as web apps, dashboards, mobile screens, and landing pages. OpenAI’s GPT-6 Astra posted the top average score at 82.7, but DeepSeek’s new V4.1 Flash came close at 81.2, or about 98% of Astra’s score. The gap in price was far wider: GPT-6 Astra cost $1.61 per finished design, while V4.1 Flash cost $0.023, roughly 1.4% as much. The benchmark also measured speed and delivery readiness. V4.1 Flash completed jobs in 5.3 minutes on average, compared with 11.1 minutes for GPT-6 Astra and 12.8 minutes for Claude Fable 5.1. Delivery rates were 57.7% for V4.1 Flash, 60% for GPT-6 Astra, and 56.7% for Claude Fable 5.1. OpenDesign said the benchmark is built to test reliable, everyday design output rather than general reasoning or coding skill, because outputs that fail to render as working webpages receive zero and are not rerun. DeepSeek’s technical report said V4.1 Flash has 552 billion total parameters but activates 8 billion to read prompts and 16 billion to generate responses, a structure it calls a Causal Encoder-Decoder design.

OpenDesign, the company behind the benchmark site OpenDesign Arena, put 13 AI models through the same batch of design tasks this week. OpenAI’s GPT-6 Astra finished first, but DeepSeek’s newest model, V4.1 Flash, came close enough to draw attention: it reached 98% of the top score while charging about 1.4% of the top price.

DeepSeek’s V4.1 Flash nearly matches GPT-6 Astra in design benchmark at roughly 1.4% of the cost 2

How OpenDesign Arena scored the models

OpenDesign Arena measures performance on everyday design work, including web apps, dashboards, mobile screens, and landing pages, on a 100-point scale.

Thirty points are assigned to whether the output actually satisfies the brief. The remaining 70 points grade design quality, including layout, hierarchy, color, and style fit.

The benchmark is built around a narrower question than most AI leaderboards ask: which model a working web designer could actually use tomorrow.

GPT-6 Astra led on score, while DeepSeek cut cost and time

On that scale, GPT-6 Astra posted an average score of 82.7. It took 11.1 minutes and cost $1.61 per finished design.

DeepSeek V4.1 Flash scored 81.2, finished the job in 5.3 minutes, and cost $0.023. Claude Fable 5.1 scored 80.3, took 12.8 minutes, and cost $3.66.

That left V4.1 Flash only 1.5 points behind GPT-6 Astra. The price gap was much wider, and the DeepSeek model also completed tasks faster.

Every other model OpenDesign tested, including Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash, and Gemini 3.8 Flash, scored lower than DeepSeek V4.1 Flash and cost more to run. That applied to 11 of the 13 models tested. Only GPT-6 Astra beat it outright, and only by a point and a half.

DeepSeek’s report outlined where the savings came from

DeepSeek’s technical report for V4.1 Flash said the model has 552 billion parameters in total, the internal settings tuned during training to store what the system has learned.

But it activates only 8 billion of them to read an incoming prompt and 16 billion to write the response. DeepSeek calls that a Causal Encoder-Decoder design, and said the same architecture supports the model’s fast completion times.

Part of a broader push by DeepSeek

This was not DeepSeek’s first attempt to narrow a capability gap at a lower price. Weeks earlier, the company’s V4 Pro model landed within 5% of Claude Fable 5 on a separate benchmark comparison while charging a fraction of Fable’s rate.

DeepSeek has also been recruiting engineers in Beijing to build its own Code Harness, with the aim of owning the full agentic stack rather than only supplying the model underneath it.

What the benchmark can and cannot show

OpenDesign’s test setup also limits what these numbers can prove. A model’s output is scored only if it renders as a working webpage in the first place. Anything blank, broken, or cut off gets a zero and is not retested.

That means the benchmark is measuring reliable, everyday design output, not general reasoning or coding skill.

GPT-6 Astra’s release and delivery rates

GPT-6 Astra, which OpenAI released on September 3, had already built a reputation for handling a wide range of tasks, from laying out a circuit board to drafting a tax return and building a 3D scene. Early testers, however, flagged it as a weaker writer than the model it replaced.

Its price and speed on OpenDesign’s chart fit that same generalist profile: slower and more expensive than DeepSeek’s cheaper entry, but still the highest scorer in the field.

DeepSeek V4.1 Flash posted a delivery rate of 57.7%, defined as the share of outputs OpenDesign judged ready to hand off without revision. GPT-6 Astra’s delivery rate was 60%. Claude Fable 5.1 came in at 56.7%.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.