Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems

N
News Editor
2026-08-23 02:15:10
An anonymous model called Ox Alpha has quickly become a focal point in the AI community after appearing on OpenRouter with a 1 million-token context window, multimodal input support for text, images, and video, tool use, and free access for now. What pushed it into the spotlight was not its listing, but its coding performance. Developer Ben Davis tested the model on 10 DeepSWE tasks and reported that it solved eight, for an 80% pass rate. In the comparison he shared, Fable 5 Max scored 65%, GLM-5.3 Max and Grok 4.6 xhigh each scored 62%, and GPT-5.6 Sol Max came in at 52%. A later run by other developers on a different DeepSWE subset produced a result of about 63%, which left Ox Alpha’s exact standing unresolved because the task sets and runtime configurations were not identical. At the same time, speculation about the model’s identity has centered on Zhipu. Analysts pointed to matching video-token behavior with GLM-5V-Turbo, a consistent 75-token gap versus GLM-5.3 across 25 prompts, and other product traits that resemble GLM routing and agent behavior. Ben Davis said he was 99% sure the model was GLM-5.x, but neither OpenRouter nor Zhipu had publicly responded as of publication. Separate debate has also formed around another anonymous model, korrine, now being tested on Code Arena.

An anonymous model named Ox Alpha has surfaced on OpenRouter and quickly turned into one of the most talked-about AI releases of the month. Chinese users have started calling it 「Niu Lai」, a nickname built on the word “Ox.”

According to information shown on OpenRouter, Ox Alpha offers a 1 million-token context window, accepts text, image, and video input, supports tool use, and is currently free to use.

Developers did not spend long wondering what it was before putting it to work. Soon after launch, the model was wired into coding agents and tested inside real code repositories. That is where most of the attention came from.

Early coding results put Ox Alpha close to leading models

Developer Ben Davis said he pulled 10 tasks from DeepSWE to test Ox Alpha. The model completed eight of them, giving it an 80% pass rate. In the comparison he posted, Fable 5 Max scored 65%, GLM-5.3 Max and Grok 4.6 xhigh each posted 62%, and GPT-5.6 Sol Max came in at 52%.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 3

DeepSWE is designed to measure real software engineering ability rather than one-shot coding performance. A model has to read a repository, identify the issue, modify code, run tests, and continue fixing errors based on the results. That makes it much closer to the kind of work coding agents are asked to do in practice.

The first result came from only 10 tasks, so the sample was limited. Other developers later ran Ox Alpha on a different DeepSWE subset and reported a result of about 63%. Because the two tests did not use the same task range or runtime configuration, the available data is not enough to pin down the model’s exact ranking.

Even with that caveat, the early runs suggest that the anonymous system has strong long-horizon coding potential and is operating near the current top tier.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 4

The leading theory points to Zhipu’s GLM line

The most widely circulated claim is that Ox Alpha may be an unreleased GLM-5.3 Flash model from Zhipu, or a multimodal version of GLM-5.3.

One blog post gathered the main evidence behind that view.

  1. The strongest point, according to the analysis, comes from the video encoder. Across four videos with different frame rates, durations, and resolutions, Ox Alpha consumed the exact same number of visual tokens as GLM-5V-Turbo. MiMo, Qwen, and GLM-4.6V reportedly showed clearly different results.
  2. The text tokenizer was also described as a close match. In tests covering 25 prompt groups, Ox Alpha and GLM-5.3 reportedly maintained a fixed gap of 75 tokens.
  3. Other traits were also said to align with Zhipu’s models. Ox Alpha refused to process audio, matching the routing behavior of GLM-5V. Its answer style, agent execution step count, and reasoning interface were also described as similar to GLM. The same blog noted that Zhipu had previously used Pony Alpha to anonymously test GLM-5.

The analysis was published at: https://ox-alpha-evidence-production.up.railway.app/

Some users also said they found small clues through direct conversations with the model. Ben Davis wrote that he was 99% sure it was GLM-5.x.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 5

That still falls short of confirmation. As of now, neither OpenRouter nor Zhipu has issued a public response.

Another anonymous model, korrine, is also under scrutiny

While Ox Alpha has dominated much of the discussion, another anonymous model called korrine has appeared on Code Arena and added to the guessing game.

One early theory linked korrine to Kimi K3.1. The reasoning came from an earlier round of speculation before Kimi K3 launched, when some people believed it had been tested under the codename kivine. Because kivine and korrine look structurally similar, the comparison spread quickly.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 6

The person who first circulated the rumor later added that a previously discussed Moonshot model might actually correspond to another codename, adamant-ananke. Under that reading, korrine could instead come from Qwen or another Chinese team.

Some commenters also pointed to MiMo V3. For now, korrine’s identity looks even less certain than Ox Alpha’s.

Why anonymous testing is becoming common

Anonymous testing is becoming a more important step before formal model launches.

On Arena-style platforms, hiding the company name and model label can reduce brand bias. Users do not know what they are selecting, so comparisons are driven more directly by output quality. The match results that accumulate in that setting can get closer to actual user experience.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 7

OpenRouter offers a different kind of proving ground. Developers can plug a model into coding agents, place it inside real repositories, and let it call tools repeatedly across software engineering tasks that can run for hours. In that setup, context stability, tool reliability, and the tendency to loop or drift during long tasks can show up quickly.

For model builders, that works like a public stress test. Teams can inspect failure cases early, check how well inference services hold up, and build real-world reputation before a formal release.

The guessing itself has also become part of the promotion cycle. Leaving a model’s identity unresolved can keep the conversation going longer.

Anonymous model Ox Alpha draws attention after coding tests place it near top-tier systems 8

No official confirmation yet

Based on the public tests so far, Ox Alpha has earned attention mainly through coding performance. If it does turn out to be a Flash model, the reported results have already raised expectations for whatever full version may follow.

That said, the identity question remains open until OpenRouter or Zhipu addresses it directly. The same is true for korrine, where the evidence is thinner and the debate is still moving.

Reference links

  • https://x.com/Adidotdev/status/2090833298713096241
  • https://x.com/davis7/status/2090669483740279155?s=20
  • https://x.com/davis7/status/2090655207831298095?s=20
  • https://x.com/MaxForAI/status/2090783750217162788

The original article was credited to the WeChat public account 「机器之心」 with the ID almosthuman2014. The listed author name was 「关注AI的」.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
180

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.