SpaceXAI launched Grok 4.7 on Sept. 21 and said it is the company’s strongest model so far for coding and knowledge work. The company said the new release keeps the same pricing and speed level as Grok 4.6, and is now available through Cursor, Grok Build, and the Grok API.
Same pricing as Grok 4.6
In its official note, SpaceXAI said Grok 4.7 uses a new base model that is larger than the one behind Grok 4.6. The company said reinforcement learning training ran longer, the task mix was made harder, and the model was aimed at work that can take several hours to finish.
SpaceXAI also said the model is better at checking its own work, handling longer context, and operating within the framework used by its resident agent, Grok Bot.
Pricing is set at $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. A faster version is also available, with double the output speed and double the price.
SpaceXAI is Elon Musk’s AI division. According to the report, xAI was folded into SpaceX in February this year, and Cursor was acquired in June. Chain News had previously outlined that history.
Company benchmarks show mixed results
SpaceXAI published a comparison table in which Grok 4.7 used the xHigh reasoning setting, GPT-5.6 Sol and Fable 5.1 used Max, and Grok 4.6 used High.
| Benchmark / Price | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol | Fable 5.1 |
|---|---|---|---|---|
| Input / output price (per million tokens) | $2 / $6 | $2 / $6 | $4 / $20 | $10 / $50 |
| CursorBench 4.0 (software engineering) | 46.3% | 40.4% | 41.7% | 51.8% |
| Terminal-Bench 4.0 (long-duration terminal work) | 38.0% | 20.3% | 37.3% | 57.9% |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% |
| Harvey legal agent benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional (clinical reasoning) | 56.7% | 48.5% | 60.5% | 62.1% |
Based on SpaceXAI’s own figures, Grok 4.7 led in electrical engineering and legal testing. It still trailed Fable 5.1 in software engineering and long-duration terminal work, and scored below the other two models in clinical reasoning. The report noted that all of these results were published by SpaceXAI itself.
Safety claims and red-team access
On safety, SpaceXAI said Grok 4.7 uses a new protection system. In the company’s own HackerBench v0.3 evaluation, it allowed only 3.3% of risky dual-use cybersecurity requests. In the LatchBio biosafety benchmark, it scored 62.4%, the highest result listed in the report.
The company also said selected cybersecurity partners have started using Grok 4.7’s red-team testing capability on an invitation-only basis for defensive research.
StarCraft: Brood War benchmark put Grok 4.6 near the bottom
A day before the Grok 4.7 release, a separate evaluation in which AI agents played StarCraft: Brood War circulated on X. The benchmark, Brood War Bench, was published by developer Ben Swerdlow and reposted on Sept. 20 by OpenClaw developer Peter Steinberger, drawing attention online.
Swerdlow built a version of StarCraft: Brood War that could only be controlled through AI agents. He then ran head-to-head matches among 19 model and reasoning-strength configurations from Codex, Claude, and Grok, for a total of 171 games.
The results showed Codex Astra at the highest reasoning setting going 18-0. Claude Fable finished with 15 wins and 3 losses, placing third, while Claude Opus 5 posted 12 wins and 6 losses. Grok 4.6’s three settings lost to every Codex and Claude setting, beating only Claude Haiku or each other. Its best-performing setup won just 2 of 18 games.
Long reasoning, very few commands
Swerdlow wrote that Grok 4.6 often spent a long time reasoning while issuing very few actual commands. In one match, the highest reasoning setting used 11,138 reasoning tokens and sent only six batches of commands over 43 minutes, without ever deploying a combat unit.
He wrote in the report that the Grok model was not smart enough to play Brood War. Older models, he said, often treated a real-time strategy game as if it were turn-based, and were destroyed while still thinking.
Swerdlow also said none of the participating models were above beginner level. Even Codex Astra and Claude Fable could not build complex armies or hold off simple attacks, and a beginner player using a photon cannon rush could beat every model in the benchmark.
The benchmark tested Grok 4.6. Grok 4.7 does not yet have a result on the leaderboard.

