OpenDesign, the company behind the benchmark site OpenDesign Arena, put 13 AI models through the same batch of design tasks this week. OpenAI’s GPT-6 Astra finished first, but DeepSeek’s newest model, V4.1 Flash, came close enough to draw attention: it reached 98% of the top score while charging about 1.4% of the top price.

How OpenDesign Arena scored the models
OpenDesign Arena measures performance on everyday design work, including web apps, dashboards, mobile screens, and landing pages, on a 100-point scale.
Thirty points are assigned to whether the output actually satisfies the brief. The remaining 70 points grade design quality, including layout, hierarchy, color, and style fit.
The benchmark is built around a narrower question than most AI leaderboards ask: which model a working web designer could actually use tomorrow.
GPT-6 Astra led on score, while DeepSeek cut cost and time
On that scale, GPT-6 Astra posted an average score of 82.7. It took 11.1 minutes and cost $1.61 per finished design.
DeepSeek V4.1 Flash scored 81.2, finished the job in 5.3 minutes, and cost $0.023. Claude Fable 5.1 scored 80.3, took 12.8 minutes, and cost $3.66.
That left V4.1 Flash only 1.5 points behind GPT-6 Astra. The price gap was much wider, and the DeepSeek model also completed tasks faster.
Every other model OpenDesign tested, including Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash, and Gemini 3.8 Flash, scored lower than DeepSeek V4.1 Flash and cost more to run. That applied to 11 of the 13 models tested. Only GPT-6 Astra beat it outright, and only by a point and a half.
DeepSeek’s report outlined where the savings came from
DeepSeek’s technical report for V4.1 Flash said the model has 552 billion parameters in total, the internal settings tuned during training to store what the system has learned.
But it activates only 8 billion of them to read an incoming prompt and 16 billion to write the response. DeepSeek calls that a Causal Encoder-Decoder design, and said the same architecture supports the model’s fast completion times.
Part of a broader push by DeepSeek
This was not DeepSeek’s first attempt to narrow a capability gap at a lower price. Weeks earlier, the company’s V4 Pro model landed within 5% of Claude Fable 5 on a separate benchmark comparison while charging a fraction of Fable’s rate.
DeepSeek has also been recruiting engineers in Beijing to build its own Code Harness, with the aim of owning the full agentic stack rather than only supplying the model underneath it.
What the benchmark can and cannot show
OpenDesign’s test setup also limits what these numbers can prove. A model’s output is scored only if it renders as a working webpage in the first place. Anything blank, broken, or cut off gets a zero and is not retested.
That means the benchmark is measuring reliable, everyday design output, not general reasoning or coding skill.
GPT-6 Astra’s release and delivery rates
GPT-6 Astra, which OpenAI released on September 3, had already built a reputation for handling a wide range of tasks, from laying out a circuit board to drafting a tax return and building a 3D scene. Early testers, however, flagged it as a weaker writer than the model it replaced.
Its price and speed on OpenDesign’s chart fit that same generalist profile: slower and more expensive than DeepSeek’s cheaper entry, but still the highest scorer in the field.
DeepSeek V4.1 Flash posted a delivery rate of 57.7%, defined as the share of outputs OpenDesign judged ready to hand off without revision. GPT-6 Astra’s delivery rate was 60%. Claude Fable 5.1 came in at 56.7%.


