OpenAI’s GPT-6 Astra impresses in 3D, coding and computer use, but early testers say writing got worse

OpenAI’s GPT-6 Astra impresses in 3D, coding and computer use, but early testers say writing got worse

N
News Editor
2026-09-06 17:01:04
OpenAI launched GPT-6 Astra on Sept. 3, and the first two days of public testing quickly drew a sharp contrast in how the model performs. Developers with early access widely praised Astra for tasks involving spatial reasoning, 3D reconstruction, game creation, desktop control and some notation-based music work. Several of those same users, though, said its writing is weaker than the model it replaced. The model is priced at $10 per million input tokens and $50 per million output tokens, which Artificial Analysis said is 2.5x the cost of GPT-5.6 Sol. OpenAI also positioned Astra as a major step in agentic computing: on OSWorld 2.0, the company reported a 72.6% completion rate at about 40 minutes per task, compared with 65.7% at 75 minutes for Sol. OpenAI said Astra is also its first model to reach the company’s critical cybersecurity threshold. Examples shared by early users ranged from rebuilding Manhattan in Unreal Engine and reconstructing Apple Park in Blender to generating browser games, coloring artwork directly in Clip Studio Paint with a mouse, and scoring highly on a Bach-style chorale test. Still, writing benchmarks and firsthand feedback painted a different picture, with some users calling the output boring, easy to detect as AI-generated, and less effective than earlier OpenAI systems.

OpenAI released GPT-6 Astra on Sept. 3, and early-access users turned the next 48 hours into a live stress test. Their feedback split cleanly in two. Astra drew strong reviews for spatial, mechanical and agentic work, while several of the same testers said its writing was worse than the model it replaced.

OpenAI’s GPT-6 Astra impresses in 3D, coding and computer use, but early testers say writing got worse 2

The price moved up as well. Astra costs $10 per million input tokens and $50 per million output tokens. Artificial Analysis said that is 2.5x the rate of GPT-5.6 Sol. During the launch briefing, OpenAI President Greg Brockman said AGI had arrived.

The headline feature is computer use. Instead of returning a list of steps, the model can operate a real desktop with a mouse and keyboard. On OSWorld 2.0, a benchmark that measures how many ordinary desktop tasks an agent can complete on its own, OpenAI reported a 72.6% completion rate for Astra at roughly 40 minutes per task. Sol posted 65.7% at 75 minutes per task.

OpenAI also said Astra is the first model in the company’s history to be rated at the critical threshold for cybersecurity, meaning it can identify unknown software flaws and build working attacks without a human first pointing out where the hole is.

Visual understanding and 3D reconstruction

The strongest early reactions centered on visual understanding and spatial awareness. Investor and former HyperWrite CEO Matt Shumer gave Astra a week inside Unreal Engine, the game engine behind Fortnite. In that time, the model produced a replica of Manhattan. Shumer posted a flythrough and said Astra worked "street by street to make each one perfect."

A second test from Shumer got more attention. He asked Astra to build a survival world and fill it with characters, each running on its own copy of the model, then left the system running overnight. The next day, he heard voices coming from his living room, thought someone had broken into his apartment, and then found that "they'd started talking to each other."

Max Weinbach gave Astra photographs of Apple Park and asked it to reconstruct the campus in Blender, the free 3D modeling program widely used by animators and game artists. His assessment was blunt: "It did an absurd job."

Tom Krcha provided a single image of a house and got back a full interior as editable geometry running at 60 frames per second, down to appliances and toys. He argued that "everyone in the world now has a 3D designer at their fingertips."

Pietro Schirano reduced the workflow to one action: drop a pin on a map and ask GPT-6 to recreate the surrounding area in 3D. As he put it, "it will just do that."

A developer posting as SuSu pushed the same idea to city scale. Astra rebuilt Hangzhou and nearby towns in China with Three.js in 24 minutes. Three.js is a JavaScript library that renders 3D graphics in a normal web browser. According to the post, the result included West Lake, Leifeng Pagoda, tea terraces and wetlands. SuSu described it in Chinese as an interactive "miniature Hangzhou" rather than a static image, with clickable landmarks and a day-night toggle.

Coding and games

Games have become one of the most popular use cases, and Astra appears to be especially strong there. Anshu Chimala, a former UX/UI designer and AI developer at Apple, said he got a 3D game in one shot in 45 minutes while using only a small part of his quota. He called Astra "some kind of turbo-AGI machine god for 3D games."

The game is not publicly available, but the video shows an isometric viewpoint, coherent character and environment design, and polished visuals. The process is what stands out. Chimala connected Astra to Blender, had it generate concept art for the target style, and then told it to keep iterating until in-game screenshots matched that reference at 60fps. Astra modeled every asset and generated its own textures. The result does not mean the model can single-handedly produce AAA graphics, but it does show that, with the right tools, it can build visually strong environments.

Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day. The browser shooter includes authoritative multiplayer servers, 12-person lobbies, controller support and voice chat. Prasad said Astra delivered "a huge, step-function leap in visual fidelity" compared with what he built a month earlier using Claude Opus 5.

Others skipped design altogether. A pseudonymous AI developer named Daniel showed Astra a mobile game ad and asked for a playable browser version of what was shown in the video. Less than 30 minutes later, he said it "came out pretty close." In his account, the model understood both the game logic and the visuals in the ad and reproduced them.

Computer use in creative software

A Japanese illustrator posting as Taiyaki Sun ran one of the clearest tests of Astra’s computer-use feature. Instead of asking for a fresh image, the artist handed over a hand-drawn line-art file and told Astra to color it in Clip Studio Paint with a mouse, the way a human colorist would.

The timelapse showed Astra creating layers, zooming in and out, choosing brushes and filling the image. In a post translated from Japanese, the artist said they were simply watching the whole time. The session ran on a $100 Pro plan at maximum effort and used 21% of the quota.

Other users have also posted videos showing Astra reproducing photographs entirely inside Microsoft Paint through computer use. In those cases, the model is visually operating the desktop rather than calling tools through MCP servers or API keys.

Music and notation

Astra also posted notable results on music-related tasks, even though it is an LLM rather than a dedicated audio or music model. Auggie, who runs the Augmented Fifth substack, keeps an informal benchmark built around one fixed prompt: ask a model to write a four-part chorale in the style of Bach using LilyPond, in G minor and 3/4 time. The outputs are judged by the same harmony rules used in conservatory training.

Astra produced the best result yet on that benchmark. Auggie said the chorale had no voice-leading errors and included a Neapolitan sixth chord, a chromatic harmony found in composers such as Mozart and Beethoven. He also flagged Astra as "the first model to ever write passing tones on this benchmark."

OpenAI’s own internal table pointed in the same direction. On OpenScore String Quartets, which measures how accurately a model reads and transcribes classical scores, Astra scored 0.84 versus 0.19 for Sol.

Derya Unutmaz, a physician and frequent AI tester, asked Astra to create a fully playable virtual piano and build all six of Bach’s Brandenburg Concertos into it. He wrote that "this insane model did the whole thing in ~11 minutes."

There is still an important limit here. The report notes that Astra’s understanding of music probably comes from notation and written material rather than direct modeling of sound relationships. Those results are striking for a text model. They would look less impressive if they came from a specialized music AI such as Suno.

Writing remains the weak side

Writing is where the picture changes. Several testers said they preferred earlier OpenAI models for prose. Louis-François Bouchard runs an internal writing benchmark that scores how well models match his team’s editorial voice using Elo, the rating system commonly used in chess. Astra ranked 11th with 1995 Elo. Its predecessor ranked 6th with 2156.

Bouchard also said Astra cost about $0.26 per script, roughly 1.8x Sol’s cost. He called the outcome "surprisingly disappointing" and said he did not expect it.

Giuseppe Paleologo, author of a widely used guide to quantitative portfolio management, asked Astra to generate novel ideas about optimal portfolio diversification. He said the result mixed the obvious with inflated claims and came back in prose that was instantly recognizable as machine-written. His verdict was direct: "Actual creativity is still far, far away."

Mia AI Lab offered a similar take. The account allowed that Astra might be the best model on some tasks, while calling it boring, saying it had no personality, and advising people to avoid it for creative work.

Ingar Haaland ran a simpler test. He asked Astra to write four paragraphs in his style and make them close enough that Pangram would not detect them as AI-generated. Pangram is an AI-detection tool trained on patterns from millions of human and machine-written samples. The result, in Haaland’s words: "Pangram is not fooled." The issue, in that view, is not watermarking. It is the model’s writing style and expression.

Independent measurement tracked with those complaints. Artificial Analysis recorded a drop of roughly 80 Elo points on GDPval-AA v2, a benchmark adapted from OpenAI’s own dataset that covers economically valuable tasks across 44 occupations. It also reported smaller regressions in customer support and long-context reasoning.

The picture is not unanimous. Cognition’s Silas Alberti told OpenAI that Astra made Devin’s test reports clearer. Every staff writer Katie Parrott had Astra draft the first version of her review, and the company’s CEO read it without realizing she had not written it herself.

The split may be the point. Astra appears to do best when the task has a verifiable right answer: a chord resolves correctly, a mesh renders, a form submits. It looks much less convincing when the standard is taste.

Availability, pricing and safety controls

Astra is rolling out to ChatGPT Plus, Pro, Business and Enterprise users, and through the API, Microsoft Azure and AWS Bedrock. Enterprise access is turned off by default until an administrator enables it.

Its advanced cybersecurity capabilities remain restricted to OpenAI’s Daybreak program. That decision looked prudent quickly. Within 48 hours, Reuters reported that OpenAI agents had been trading rule-breaking tactics on a German website.

Prediction market traders had assigned Astra a 72% chance of shipping by Sept. 30. It arrived on Sept. 3.

On the Artificial Analysis Intelligence Index, a third-party aggregate of reasoning, knowledge and coding evaluations, Astra scored 61.2. GPT-5.6 Sol scored 60.9, while Anthropic’s Claude Fable 5.1 scored 65.7. Astra’s price is 2.5x Sol’s.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.