ARTICLE•NEWS 2+

Ox Alpha The Mysterious AI Model Shaking Up OpenRouter

23 August 2026•
Image for Ox Alpha The Mysterious AI Model Shaking Up OpenRouter

On August 20, 2026, a new artificial intelligence model simply appeared on OpenRouter with no lab name, no press release, and no model card. The only word listed in the provider column was “Stealth”. Within hours, the developer and AI research community around the world instantly turned into amateur detectives. They tried to solve one simple yet intriguing question, namely who actually built this model called Ox Alpha.

This phenomenon is not the first of its kind. However, the way it spread and the intensity of speculation that accompanied it made Ox Alpha one of the hottest topics in the AI world in late August 2026.

A Debut Without an Identity

Ox Alpha was first spotted on OpenRouter under the model ID stealth/ox-alpha. The technical specifications shown on the model’s listing page immediately caught the market’s attention.

  • A context window capacity of up to 1,048,576 tokens, roughly equivalent to 1 million tokens.
  • Completion support capable of handling up to 131,072 tokens in a single process.
  • The ability to process text, images, and video input. Video support is rare because only a handful of frontier-class models can handle it well.

What made this model even more attention-grabbing was its pricing policy. The service is completely free, both for prompt and completion tokens, during a preview period expected to last about one week. The operator behind Ox Alpha even claims to have service capacity of up to 100 trillion tokens per day. This figure indicates that some serious heavy-duty computing infrastructure is operating behind it.

The model is positioned as a reasoning model designed specifically for coding, long-horizon agentic work, and production use. OpenRouter itself explicitly states that it acts only as a request router. It is not the developer, owner, or actual provider of the model. In other words, behind the name “Stealth” sits an anonymous third party that deliberately chose not to reveal its identity during this trial period.

Not long after its appearance, the OpenCode coding agent platform also announced support for Ox Alpha. They offered free access with a one-million-token context, multimodal support, a zero data retention policy, and very loose rate limits. This combination prompted thousands of developers to try it right away on real projects.

Surprising Benchmark Scores

The story went even more viral thanks to unofficial testing done by a developer named Ben Davis. He ran Ox Alpha through ten tasks from DeepSWE, a benchmark designed to test long-horizon software engineering capability.

The result was stunning because Ox Alpha scored around 80 percent. That number could even be higher after correcting for one small mistake related to handling the character “x”. For comparison, Claude Fable 5 recorded a score of 65 percent, while GPT-5.6 Sol only reached 52 percent on the same set of tasks.

These numbers immediately triggered a wave of discussion. Some parties considered this strong evidence that Ox Alpha is a frontier-class model that has never been officially announced. However, many others remained skeptical.

Critics pointed out that the DeepSWE testing harness only supports bash, so its coverage is quite limited. A sample size of just ten tasks was also considered too small to draw solid conclusions, since variance can be very large. Furthermore, there were reports that Ox Alpha actually lost to smaller models, such as Muse Spark, on simpler real-world tasks. This raised suspicions of possible benchmaxxing, the practice of optimizing a model specifically to perform well on certain benchmarks without genuinely reflecting its general capabilities.

The Identity Hunt and Why GLM Became the Strongest Candidate

When a lab chooses not to reveal its identity, the AI research community usually turns to a digital forensics method called fingerprinting. This method analyzes technical traces unintentionally left behind by a model. In the case of Ox Alpha, several types of fingerprinting tests were carried out in parallel by different parties, and their results surprisingly pointed to the same conclusion.

One of the most thorough analyses was done by Ben Davis himself. Besides running performance benchmarks, he also conducted video encoder tests, tokenizer analysis, response style checks, and even the model’s refusal behavior toward audio input. He stated he was fairly confident that Ox Alpha is a model from the GLM-5.x family because all of those elements showed a match.

Another equally fascinating analysis came from a researcher with the account @unclecode. He built a kind of dedicated forensics tool consisting of nine infrastructure probes. The tool covered tokenizer tests, API error codes, and hidden templates.

That tool was then run against twelve suspect models at once. The result was that only one model family consistently matched across all tokenizer tests, which was GLM. This researcher’s conclusion was brief but firm, stating that technical infrastructure traces cannot lie.

Another clue came from infrastructure performance characteristics. Researcher Tim Dettmers noted that Zhipu, the company behind GLM, is known as one of the few labs that ships models with relatively slow partial-prefill speeds but fast output speeds.

This unique pattern was also observed in Ox Alpha. There is an additional note that Ox Alpha runs about 30 percent faster than GLM 5.3, indicating a possibly newer version or further optimization.

Another frequently cited piece of evidence is a highly precise tokenizer match. One analysis reported that Ox Alpha’s tokenization pattern differs from GLM 5.3 by a consistent margin of only 75 tokens. This is compounded by similar speeds and cache hit rates. These findings strengthen the suspicion that Ox Alpha may be a new multimodal variant of the GLM family, most likely called GLM 5.3 Vision.

Some reports went even bolder in stating their confidence level. One independent analysis cited a probability of 90 to 99 percent that Ox Alpha is an unreleased GLM-5.x model. This conclusion was based on a combination of matching visual encoders, token consumption patterns, response rhythm, and agentic execution steps nearly identical to GLM-5.3V and other related multimodal variants.

Alternative Theories Circulating

Although the GLM suspicion appears most dominant and is backed by relatively strong technical evidence, that does not mean other theories are not circulating in the community.

  • Xiaomi MiMo

A number of users suggested that Ox Alpha might actually be the work of Xiaomi’s MiMo team. That lab does have a track record of releasing models in stealth mode in the past, one of which used the codename “Hunter Alpha”. However, this theory was considered weak after further testing revealed differences in audio endpoint behavior and mismatched tokenizer patterns.

  • Google Gemini

There is speculation linking Ox Alpha to the Google Gemini team. This suspicion arose because several users reported that the model occasionally referred to itself as a Gemini variant when asked directly. However, such self-identification claims are considered unreliable. Language models do hallucinate about their own identities quite often, especially after going through fine-tuning or distillation from other models’ data.

  • Other Candidates

Other names mentioned in public discussions include DeepSeek, Qwen, Kimi, and speculation that Ox Alpha is a product called “Composer” built on top of Kimi or GLM foundations. However, none of those theories is supported by technical fingerprinting evidence as strong as the case for GLM.

Why This Pattern Feels Familiar

For anyone who has followed the AI industry over the past few months, the pattern behind Ox Alpha is nothing new. Ox Alpha is the fifth stealth model to appear on OpenRouter within roughly the last six months.

The four previous anonymous releases also sparked similar speculation. In the end, however, all of those models were officially claimed by laboratories from China.

  • GLM-5 by Zhipu AI
  • MiMo-V2-Pro by Xiaomi
  • Ling-2.6-flash by Ant Group
  • LongCat-2.0 by Meituan

Interestingly, this stealth launch pattern has been openly acknowledged before. In the GLM-5 technical report, the Zhipu team included an easter egg confirming that an experiment named “Pony Alpha” was indeed GLM-5 released anonymously on OpenRouter. The strategy was used to gather purely unbiased user feedback without the bias of a big brand name.

In that report, they mentioned that at the time, most users guessed the model was Claude Sonnet 5, DeepSeek, or Grok, before its true identity was finally revealed.

Given this precedent, it is understandable that many believe the stealth launch strategy has become a favorite marketing and testing tactic among Chinese AI labs. It is an effective way to test a model’s performance in the real market before launching officially with full branding.

What to Consider Before Trying It

Despite the tempting free access to a model with a one-million-token context window, there are important things to consider before using Ox Alpha for serious work. The model’s OpenRouter page states that the third-party provider stores prompts and completions, although the data is claimed not to be used for retraining the model. Other usage terms are governed by what are called the Stealth Model Terms.

There is also a personal note worth sharing from trying it firsthand. Even though the quota is nearly unlimited and free, Ox Alpha’s responses feel noticeably slow compared to other models, whether free or paid, such as HY3, DeepSeek, and a number of other popular models. On top of that, generation sessions occasionally fail or stop midway, so requests have to be sent again. For casual experimentation or tasks that do not demand high speed, this is still tolerable. However, if you intend to rely on it for work with tight deadlines or need consistent output, this reliability factor still deserves consideration, especially since the free period itself only lasts briefly, around August 20 to 27, 2026.

Since the true identity of the operator remains unknown, most observers advise that Ox Alpha should only be used for exploration, open-source projects, or repositories already sanitized of sensitive data. For work involving proprietary code or confidential company data, users are advised to wait until there is an officially identified operator with clear and verifiable service terms.

Conclusion

As of now, no laboratory has officially claimed to be behind Ox Alpha. However, based on the accumulated evidence from various technical fingerprinting methods, the strongest suspicion still points to Zhipu AI’s GLM family. The indicators show up in the tokenizer match, video encoder patterns, processing speed, and response style. The model is likely a new multimodal variant such as GLM-5.3 Vision, or perhaps an even newer generation.

Will this suspicion prove correct like the previous patterns, or will Ox Alpha instead become the exception that breaks community expectations? Only time will tell. What is certain is that the Ox Alpha phenomenon once again demonstrates how sophisticated the global AI community has become at conducting technical investigations. They can uncover a model’s digital trail just from how it processes tokens, answers questions, and behaves behind the scenes, all without any official confirmation from its creator.