Silicon Valley saw three major AI developments in one night. Google unveiled Gemini 3.8 Flash. Meta followed with Muse Spark 1.3. Startup Mostik, a company that had barely entered the broader conversation before, appeared in a WIRED feature.
If the first two items are viewed on their own, the sequence looks familiar: another benchmark-heavy evening in AI, with one model reaching the top of DeepSWE and another overtaking it a few hours later. Put together, the three stories suggest something else. The most important change may not be who led a ranking for a few hours, but how AI economics is being measured.
For the past few years, the standard pricing reference in AI has been cost per million tokens. With agents taking on longer and messier workflows, that unit is starting to lose some of its usefulness. The more relevant question is shifting toward the cost of fixing a bug, finishing a research task, or letting an agent work continuously for three hours.
Seen from that angle, the night’s announcements all point in the same direction from different starting points. Google is pushing frontier-level capability into cheaper models. Meta is trying to get agents to complete the same work with fewer tokens and fewer tool calls. Mostik is asking a deeper question: if both parties in a conversation are AI systems, why should they still communicate through language built for humans?
Google pushes Gemini 3.8 Flash deeper into frontier territory
Google described Gemini 3.8 Flash as its most powerful reasoning and coding model to date. It is also the third update to the Flash line in six weeks, moving from 3.6 to 3.7 and now 3.8.
Flash models were long associated with a straightforward trade-off: faster and cheaper, but not quite at the level of the strongest systems. Google is now trying to change that definition.
The main gains in Gemini 3.8 Flash are centered on long-horizon coding and agentic workflow. On DeepSWE v1.1, a benchmark aimed at long-duration software engineering, it scored about 74%. That benchmark differs from older coding tests because it does not simply give a model an algorithm problem. It places an agent inside a code repository and requires it to understand the issue, search the codebase, edit files, call tools, run tests, and keep debugging after failure. A single task can take dozens or even hundreds of steps.
On that test, Gemini 3.8 Flash briefly reached the top. Claude Opus 5 was also around 74%, GPT-5.6 Sol came in at about 73%, and the previously leading Fable 5 was around 70%. In capability terms, the gap among top systems has narrowed.
The stronger contrast appears in cost. Gemini 3.8 Flash launched at $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google said the average cost of completing a full DeepSWE task was $2.36.
For comparison, Claude Opus 5, with roughly the same score, reached an average of $11.84 per task. GPT-5.6 Sol, at about 73%, came in at $6.46. That means similar software-engineering performance can now carry task-execution costs that differ by several multiples.
That matters more in the age of agents than it did in the chatbot era. A chatbot answer that costs a few cents more or less is often invisible to most users. An agent can run 100 steps in a coding task, search dozens of websites in a research workflow, read hundreds of pages, or stay active for hours inside an enterprise setting.
Google also said 3.8 Flash will "work harder" on difficult tasks, increasing reasoning and tool use when needed. The point is not that the model has been forced to think less. It is that tokens have become cheap enough to support longer loops. In effect, Google is trying to compress capabilities that once belonged only to expensive frontier models into Flash pricing.
Meta’s Muse Spark 1.3 raises the focus on task efficiency
Roughly three and a half hours after Google’s launch, Meta re-entered the picture with Muse Spark 1.3.
Based on Meta’s published evaluation, Muse Spark 1.3 scored 75.4% on DeepSWE v1.1, ahead of Gemini 3.8 Flash, Claude Opus 5, and GPT-5.6 Sol. The jump is notable because Muse Spark 1.2 had been at about 55% on the same benchmark. From version 1.2 to 1.3, the increase was about 20 percentage points.
The article argues that the more important numbers sit elsewhere. Meta also reported about 20% fewer tool calls and about 25% lower token consumption.
Those figures connect more directly to commercial agent deployment. In many agent workflows, the expensive part is not ordinary reasoning but drift. A coding agent that misunderstands the task at step 20 may keep editing five files, run three rounds of tests, search large parts of a codebase, and only later discover that it took the wrong path. Dozens of tool calls and tens of thousands of tokens can be wasted before the mistake is corrected.
A model that identifies ambiguity earlier, asks for clarification sooner, or admits it cannot complete a step is not just behaving better. It is cutting cost.
Muse Spark 1.3 strengthened several capabilities that fit this pattern. It can maintain multiple workflows inside long threads, ask follow-up questions when a task is unclear, request help when it cannot solve a problem, confirm before irreversible actions, and retain the original constraints more reliably over long runs.
Meta also emphasized awareness of capability boundaries: knowing what the model can do and what it cannot. In the chatbot cycle, that may have looked less dramatic than a 20-point benchmark jump. In the agent cycle, it translates into money saved. Every avoided mistake is a form of inference optimization.
Viewed together, Google and Meta are competing on more than benchmark peaks. The unit of comparison is starting to move from the price of 1 million tokens to the price of completing a real-world task from start to finish.
Mostik questions whether those tokens need to exist at all
If Google and Meta are still asking how to finish tasks with cheaper or fewer tokens, Mostik is pushing into a more basic layer: why should those tokens exist in the first place?
Mostik means "bridge" in Russian. CEO Sasha Malysheva is described as the main developer of the method. The company’s chief scientist is Stanislav Smirnov, a University of Geneva professor and the 2010 Fields Medal winner.
The company is working on a difficult idea that sounds simple when stated aloud: getting two AI models to stop communicating through natural language.
Most multi-agent systems today work like this. Model A receives information, writes out a conclusion in hundreds or thousands of tokens, and Model B reads that text back in, reconstructs its own internal representation, and continues the reasoning process.
The problem is that the underlying computation inside large models is not really operating in English or Chinese. It is happening in high-dimensional continuous mathematical representations. The article compares the setup to two computers that could transfer data directly, except Computer A first prints the file into hundreds of pages and Computer B then uses a camera to OCR every page back into machine-readable form. Much of the field has focused on lowering the printing bill. Mostik wants to remove the printer.
The company is trying to build a bridge between the internal representations of different models so they can exchange latent representation directly rather than first generating natural-language messages.
According to an experiment cited by WIRED, Mostik linked the full GLM-5.2 753B with a 4B version of Qwen 3.5 that can run on mobile devices. The resulting hybrid system delivered capability between the two models while cutting inference cost to 1/20 of full GLM-5.2.
The article also notes that it is too early to say Mostik has solved model communication. Latent communication did not appear for the first time yesterday. Over the last several years, researchers have experimented with exchanging embeddings, hidden states, KV cache, and other internal forms.
The hard part is compatibility. Different models have different architectures, parameters, training data, and internal coordinate systems. Getting them to truly understand one another’s latent space is a difficult problem on its own. Smirnov said there is still no mature mathematical language for describing shared representation across models.
Even so, the 1/20 figure stands out because it opens the door to a different possible AI architecture.
A possible split between tiny local models and remote frontier systems
Over the last two years, discussion around on-device AI has mostly focused on how to fit larger models into phones: 7B, 4B, 3B, 1B, with constant work on distillation, quantization, and compression.
If Mostik’s route works, a different structure could emerge. A future phone may not need a large all-purpose model running locally. It may only need a 0.xB or a few-B model that is cheap enough to stay active all the time, track which app the user is in, what just happened on the device, and what state the hardware is in, while handling most simple and frequent tasks.
When a harder problem appears, that small local model could call a remote larger system through latent-space communication. The key difference is that it would not need to resend 100,000 tokens of context to the cloud for the remote model to reread from scratch. It might only transmit a heavily compressed latent state. After the remote model finishes a more difficult reasoning task, it may not need to return thousands of tokens of natural-language explanation either. It could send back a new internal representation directly.
In that setup, the local small model handles frequent, cheap, always-on work. The cloud frontier model handles the infrequent but difficult reasoning. A bridge between them carries information efficiently. If that becomes viable, the decline in inference cost may not stop at a 30% or 50% API price cut. It could change how AI computation is organized.
What changed in one night
Look back at the three developments and they line up around the same theme. Google released Gemini 3.8 Flash and pushed frontier-level capability into Flash pricing. Meta released Muse Spark 1.3 and showed that an agent can finish more work with fewer tokens and fewer tool calls. Mostik went one step lower and explored how to prevent some tokens from being generated at all during model-to-model communication.
Taken together, the message is that intelligence is becoming cheaper at high speed. The decline is no longer limited to lower API pricing. Model prices are falling, the amount of compute needed to finish tasks is falling, and even the way models exchange information is being redesigned.
The article ends by framing the bigger question not as who briefly held the top benchmark score, but how cheap highly capable AI can ultimately become. Many technologies do not break out when they first become usable. They break out when they become cheap enough to use freely.
Today, paying tens of dollars for an agent to complete a small everyday task may make little sense. If that drops to a few cents, many applications that currently seem uneconomic could become practical. The article adds that running dozens of agents around a single person, 24 hours a day, is unrealistic now. If inference costs fall by another one or two orders of magnitude, that could start to look normal.
That leaves a different takeaway from the night’s benchmark battles. Gemini’s 74% and Muse’s 75.4% may soon be replaced by newer numbers. The harder question is what happens once strong AI becomes cheap enough to call on at will.

