Early test results for two unreleased Claude models have started circulating, with public discussion centering on the codenames claude-marshmallow-eap and claude-melon-eap. Testers quoted in the source material said the models performed strikingly well on 3D reinforcement learning tasks and architectural spatial layout problems, while also showing unusually heavy compute usage.

One-shot outputs stood out in 3D and layout tests
According to the report, one tester said Anthropic is pushing aggressively on 3D reinforcement learning. The two models were described as capable of directly understanding and generating complex 3D coordinates and spatial relationships, with architectural layout results drawing particular attention. In several of the shared impressions, the outputs were said to be generated in one shot.
The source article framed this as notable because spatial reasoning has long been a weak spot for large language models. Tasks such as understanding the physical coordinate relationships of objects in a 3D room, accounting for gravity, or planning building clusters with complicated circulation paths and load-bearing structures have generally been difficult for LLMs. The early feedback around Marshmallow and Melon suggests those limits may have shifted, at least in the tests described.

The same report said the models also appeared strong in handling topology, geometry, and physical constraints. Testers claimed they did not need repeated prompt adjustments to fix mistakes. Instead, the models were described as doing deeper internal reasoning before producing a complete 3D scene output in a single pass.
Model names and traffic traces fueled the leak narrative
The naming convention itself drew interest. The article said the "food + EAP" pattern follows an earlier internal-style label seen in July, claude-horchata-eap.
As for where the new models surfaced, the report cited data that developers said they intercepted from API traffic logs. Between 00:45 and 00:57Z on Aug. 21, the ID claude-marshmallow-ht-eap was reportedly called 57 times. It later showed up in Claude Code’s model list under the label "Custom model" and was displayed with a 1 million-token context window.

At the time covered by the source, neither model had a first-party API endpoint. The article said that suggested access may have been limited to red-team testers or internal core personnel, though scattered early testing feedback had already spread widely enough to trigger heavy discussion.
Marshmallow was seen as slightly stronger than Melon
In the testing feedback cited by the article, Marshmallow was viewed as slightly stronger overall than Melon. For everyday conversation and logical interaction, one line of commentary held that Marshmallow even felt better than the current Opus 5, with a more natural and pleasant conversational style.
That said, the same source also noted that both models still seemed to sit below Anthropic’s top "Fable" tier. Whether they are tied to a rumored Opus 5.1, a Sonnet 5.1, or a fresh Haiku iteration remained unresolved in the reporting.

Heavy "thinking token" use became a recurring theme
Beyond the spatial reasoning discussion, testers repeatedly pointed to another shared trait: very high compute consumption. Multiple people said the two models used an unusually large amount of compute because they consumed large volumes of what the report called "thinking tokens."
One tester wrote on X: 「I noticed one thing both of these models have in common: they use massive amounts of thinking tokens, to the point that during testing I repeatedly hit the

The source explained this as a shift away from the older style of model interaction, where an AI system behaves more like a fast responder that generates an answer token by token almost immediately. In contrast, the new Claude models were described as producing a large number of "invisible" tokens before giving a final answer, using them for deeper self-reasoning, chain-of-thought construction, and logical trial and error in the background.
The article interpreted that behavior as a strong signal that Anthropic is working on deep reasoning bottlenecks. It also linked the observation to an existing industry view mentioned in the report: competition in large models may be moving from training-time compute scaling toward inference-time compute consumption.
The timing invited comparisons with Opus 5
The leak also drew attention because of when it appeared. The article noted that Anthropic released its flagship Opus 5 on July 24, 2026, less than a month before these new model discussions surfaced. Before that, Sonnet 5 had been released on June 30 and was described as broadly well received.

Opus 5, however, was portrayed in the report as having run into a rougher reaction. One developer quoted on X wrote: 「The negative feedback around Opus 5 is so severe that Anthropic had to immediately push out two new models to try to replace it... this is killing me.」
Against that backdrop, the sudden appearance of Marshmallow and Melon produced two main lines of speculation in the source article. One is that the pair could be closely tied to Opus 5.1, serving as enhanced revisions aimed at fixing shortcomings in Opus 5, with the thinking-token mechanism used to push reasoning ability higher. The other is that, given the claimed speed and cost-performance profile, they may instead belong to updated Sonnet or Haiku lines, especially since they were not labeled under the top-end "Fable" name.
No official conclusion yet
For now, the formal positioning of the two models remains unclear. The information in the source report came from developer-cited API traffic traces, model list sightings, and tester comments posted on X, rather than an official product announcement from Anthropic.

The article closed by citing one developer’s reaction: 「The next few weeks to months are going to get very interesting!」 Whether Marshmallow and Melon can change the narrative after the Opus 5 release is still an open question in the public discussion captured by the report.
Referenced materials in the source included posts on X from Lentils80 and NFT_Chen. The original article was credited to the WeChat public account "新智元," with ASI启示录 listed as author and Aeneas as editor.

