Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing

N
News Editor
2026-09-03 13:23:19
Meta has launched Muse Spark 1.3, putting fresh pressure on Google’s Gemini 3.8 Flash in both benchmark rankings and pricing. According to the source article, Meta CEO Mark Zuckerberg described the release as the company’s biggest jump so far in coding and agent work, while Meta Superintelligence Labs lead Alexandr Wang said the new version cuts unnecessary turns, reduces tool calls by about 20%, lowers token use by about 25%, and holds requirements better in long-running tasks. The model’s reported scores were a major part of the launch narrative. Muse Spark 1.3 posted 75.4 on DeepSWE v1.1, ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0. It also scored 59.4 on SWEAtlas CodeBase QnA, 88.8 on Terminal-Bench 2.1, 98.1 on MRCR 512K-1M, and 62 on the Artificial Analysis intelligence index. On cost, the article said Meta kept pricing at $1.25 per million input tokens and $4.25 per million output tokens, while a contributor tier drew attention for especially low pricing. The report also noted skepticism. BridgeMind said recent AI launches have delivered surging benchmark numbers without matching gains in real-world experience. Artificial Analysis data cited in the article also showed weaker results for Muse Spark 1.3 in AA-Omniscience, where a higher abstention rate reduced hallucinations but pointed to limits in absolute knowledge gains.

Meta has released Muse Spark 1.3, escalating the contest over high-end AI models only hours after Google rolled out Gemini 3.8 Flash.

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 2

The source article said five frontier models had already been released in the first three days of September: Anthropic’s Fable 5.1 and Mythos 5.1, Google’s Gemini 3.8 Flash and Cyber, and Meta’s Muse Spark 1.3. Zuckerberg announced the new model on X, calling it a major step for coding and agent work and saying it delivers frontier performance at a cost that is almost negligible.

Alexandr Wang, who leads Meta Superintelligence Labs, followed with several posts of his own. He said Muse Spark 1.3 reduces ineffective turns compared with Muse Spark 1.2, cuts tool calls by about 20%, lowers token consumption by about 25%, and does a better job preserving user requirements in long-running tasks. In his words, users would clearly feel the jump.

Muse Spark 1.3 climbs into the top tier on benchmarks

The article framed the launch as an immediate challenge to Gemini 3.8 Flash. On benchmark results, Muse Spark 1.3 was described as making a larger leap than the newly released Google model.

Under max reasoning intensity, Muse Spark 1.3 reached 62 on the Artificial Analysis intelligence index, placing it ahead of GPT-5.6 Sol and Fable 5 and into the current top tier. The report also noted that Meta has iterated through four Muse Spark versions in five months.

On DeepSWE v1.1, which the article described as a benchmark that better reflects real long-horizon coding ability, Muse Spark 1.3 scored 75.4. That was higher than Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, and it also put the new Meta model ahead of Gemini 3.8 Flash.

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 3

  • Muse Spark 1.3: 75.4
  • Claude Opus 5: 74.0
  • GPT-5.6 Sol: 73.0

The source article listed several other developer-focused metrics as well:

  • SWEAtlas CodeBase QnA: 59.4, above GPT-5.6 Sol and Opus 5
  • Terminal-Bench 2.1: 88.8
  • MRCR 512K-1M: 98.1

For comparison, GPT-5.6 Sol scored 73.8 on MRCR 512K-1M, according to the article. On agent-oriented tests such as JobBench and OSWorld, Muse Spark 1.3 was said to trail Opus 5 by only a few percentage points.

Artificial Analysis was a recurring reference point throughout the piece. Muse Spark 1.3 was reported to have matched Claude Fable 5 at 62 on the intelligence index. The article also said Muse’s input cost is 8 times lower than Fable 5’s and its output cost is nearly 12 times lower.

Wang used stronger language in public, according to the source. He mocked Gemini’s standing on the Artificial Analysis intelligence index and said it could only breathe other models’ exhaust.

Pricing becomes the other centerpiece of the release

Performance was only part of the pitch. Cost was the other.

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 4

The source article said Muse Spark 1.3 did not follow the old route of simply pushing scale and stacking parameters. Instead, it was trained specifically for long-horizon coding and stronger instruction following, with the aim of reducing unnecessary interaction rounds and making outputs more concise. That, the article argued, translated into cleaner code style, fewer wasted tokens, and a sharp drop in cost.

One developer test was highlighted in detail. Developer @SPAC89 compared Muse Spark 1.3 Ultra Contributor with Fable 5.1 xHigh using a prompt that included a strict self-improvement rule: if an independent judge scored the result below 9.5 out of 10, the agent had to keep improving the code and try again.

Both models ran for about two hours. During that period, Muse Spark 1.3 reportedly completed 20 self-improvement loops, launching three agents each time. The source said that amounted to more than 60 agent runs within roughly two hours, with a total cost of less than $1. The developer described the result as far beyond expectations and called it a work of art.

On headline pricing, Meta kept the model at $1.25 per million input tokens and $4.25 per million output tokens. The article added that the contributor version, muse-spark-1.3-contributor, sits near the floor of overseas model pricing, with the tradeoff that user data can be used to improve Meta’s products.

Mehul Mohan, founder of Codedamn, was quoted as saying the contributor-tier pricing looked unreal and calling it a US AI lab’s “DeepSeek pricing moment.”

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 5

Using Artificial Analysis data, the article said Muse Spark 1.3 has a per-task cost of $0.55 among models with similar intelligence levels, defined there as scores above 59. The same comparison listed:

  • Grok 4.6: $0.94
  • GPT-5.6 Sol: $0.95
  • Claude Opus 5: $1.23

Developer Kartik also said Muse Spark 1.3 appears to show a form of “loop-depth” reasoning similar to Sol and Fable. In that description, the model does not produce an answer in a single pass but performs internal logical work first, which helps explain its token efficiency.

Meta’s larger target is a 24/7 personal agent

The release was also placed in the context of Meta’s broader roadmap.

The article looked back to Muse Spark 1.1, released in July, which focused on a 1M context window and computer-use capabilities. Less than two months later, version 1.3 was described as pushing long-horizon coding performance up to the level of GPT-5.6. That speed of iteration was presented as a result of Wang taking over Meta Superintelligence Labs.

Two phrases were repeated in the source piece: “agents” and “cheap enough to ignore.” The report said Meta is effectively aiming at “agent engineering” as the next major wave.

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 6

It cited Wang’s interview with Axios, where he said usability improvements are there to serve Zuckerberg’s stated ambition of building a personal agent that stays online 24 hours a day, seven days a week.

The article broke that goal into two requirements. First, the model has to be smart and stable enough to manage very long threads, handle multiple workflows in parallel, ask clarifying questions when instructions are vague, seek human help when blocked, confirm before taking critical actions, and avoid hallucinations. Second, it has to be cheap. If a personal agent costs tens of dollars a day to run, the article argued, it will remain a niche product.

On that reading, Muse Spark 1.3 is meant to prepare the ground for always-on personal agents. The source also said Meta researcher Sun Zhiqing, described there as a Peking University alumnus, disclosed that the team conducted another round of post-training on the earlier avocado model and achieved better test scaling.

Skepticism over benchmark chasing is not going away

Not all of the reaction was celebratory.

The article quoted AI group BridgeMind as saying Gemini 3.8 Flash and Muse Spark 1.3 were released on the same day and both looked strong on benchmarks. Muse Spark 1.3 is currently tied with Fable 5 on the intelligence index, BridgeMind said, “I want to feel excited. But I don’t.”

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 7

BridgeMind’s argument was that recent AI launches have followed the same pattern: scores jump, real-world experience does not. It said trust in benchmarks is falling week by week, and trust in labs that game those benchmarks is thinner still.

The source article also mentioned a stress test by online users involving Muse Spark 1.3 and Gemini Flash 3.8z under backend-level 3D constraints. It said those results were used to argue that both models showed weaknesses, and it quoted one view that “Muse Spark 1.4 is not that good” while “Gemini Flash 3.5 is purely benchmark farming.”

That fed into a broader point in the article: many labs are now training to benchmarks. DeepSWE, Tau3-Bench, and GPQA were cited as examples of tests becoming explicit targets.

Even within Artificial Analysis’ deeper report, the article said, Muse Spark 1.3 showed signs of regression in places. On AA-Omniscience, both the xhigh and max variants declined. The stated reason was a higher abstention rate: when uncertain, the model is now more likely to withhold an answer instead of making one up. That lowers hallucinations, but it also suggests that the growth in absolute knowledge may not be as dramatic as the headline scores imply.

The next phase of model competition centers on long tasks and cost efficiency

The source closed with a set of conclusions drawn from the same-day releases of Muse Spark 1.3 and Gemini 3.8 Flash.

Meta’s Muse Spark 1.3 raises the pressure on Gemini with stronger benchmark scores and lower pricing 8

First, the benefits of scaling base models through parameter growth alone are approaching their limits. The article argued that future gains may come less from raw size and more from architectural choices, such as loop-like reasoning mechanisms and stricter cuts in tool use.

Second, the contest is moving from “who talks better” to “who gets more work done.” In that frame, the core barrier is no longer a single benchmark score. It is the real ability to complete long-horizon tasks at very low cost.

The article argued that a model capable of completing 60 trial-and-error cycles for under $1 could change business models altogether.

The references listed in the source included Meta’s research blog, Artificial Analysis, X posts from Alexandr Wang and Sun Zhiqing, and the developer test shared by @SPAC89.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.