---
title: "Mysterious AI Model Surpasses GPT-5.6 and Claude in Coding Tests; Technical Features Point to Z.AI's Unreleased New Model"
type: "News"
locale: "en"
url: "https://longbridge.com/en/news/296597174.md"
description: "An anonymous AI model named \"stealth/ox-alpha\" has appeared on OpenRouter. Independent researcher Ben Davis found that its technical features, including video encoder, tokenizer, and output style, closely resemble Z.AI's GLM series, outperforming mainstream models like GPT and Claude in certain coding benchmarks. Davis stated he is \"99% certain\" the model originates from Z.AI, though this has not yet been officially confirmed"
datetime: "2026-08-21T09:35:13.000Z"
locales:
  - [zh-CN](https://longbridge.com/zh-CN/news/296597174.md)
  - [en](https://longbridge.com/en/news/296597174.md)
  - [zh-HK](https://longbridge.com/zh-HK/news/296597174.md)
generator: "portal-rs"
---

# Mysterious AI Model Surpasses GPT-5.6 and Claude in Coding Tests; Technical Features Point to Z.AI's Unreleased New Model

An anonymous AI model has quietly emerged on the model distribution platform OpenRouter. Tests by independent researchers show that the model, named "stealth/ox-alpha," not only surpasses mainstream models like GPT and Claude in certain coding benchmarks but also exhibits technical features strongly pointing to Z.AI's unreleased next-generation multimodal flagship model.

According to a test report released by technical researcher Ben Davis on August 21, ox-alpha went live on OpenRouter on August 20 and currently offers free access for one week. The model supports text, image, and video inputs, possesses reasoning capabilities, and features a context window of 1.048 million tokens.

Davis stated on X that **he is "99% certain" that ox-alpha belongs to Z.AI's GLM-5.x series, citing multiple pieces of evidence such as the video encoder, tokenizer, output style, and audio rejection behavior. His tests also showed that ox-alpha performed better than GPT and Claude in some comparative tests.**

However, it is important to emphasize that these conclusions are based on personal testing and technical inference and have not yet been officially confirmed by Z.AI.

## Video Encoder Becomes Key Evidence

Davis's tests traced the technical origins of ox-alpha across multiple dimensions, including the video encoder, tokenizer, audio interface, and output style. Among these, **the high match in the video encoder is considered the strongest evidence.**

Tests showed that in four sets of controlled videos, ox-alpha's video token consumption was identical to that of GLM-5V-Turbo, with both exhibiting the same three characteristics: frame-rate-independent frame sampling, a duration scaling ratio of approximately 147 tokens per second, and a frame-by-frame resolution scaling mechanism.

In contrast, candidate models such as MiMo v2.5, Qwen 3.8 Max, and GLM-4.6V displayed significantly different video encoding features.

Tokenizer tests also pointed to GLM. **The report stated that after testing 25 different prompts, ox-alpha's token count matched GLM-5.3 exactly, with only a fixed +75 token hidden wrapper difference, suggesting that the two may share the same vocabulary.**

Furthermore, ox-alpha's behavior of rejecting audio inputs aligns with GLM-5V, whereas MiMo v2.5, a primary competitive candidate, supports audio input, further weakening the likelihood of it being the source.

In terms of output style, ox-alpha uses approximately 1.3 emojis per thousand characters, which is relatively close to the GLM/Qwen series; in comparison, Claude, GPT-5.6, and Grok had emoji usage rates close to zero under the same test conditions.

## Coding Capabilities Surpass GLM-5.3?

Regarding capability tests, the report cited DeepSWE benchmark data, stating that ox-alpha passed 8 out of 10 deterministic tasks, achieving a Pass@1 rate of 80%.

For comparison, Claude Fable 5 had a pass rate of 65%, while GLM-5.3 and Grok 4.6 both achieved 62%, and GPT-5.6-sol reached 52%. However, since the number of test runs varied across models and the sample size for ox-alpha is currently small, these results require verification through more independent tests.

Notably, in the "meriyah-explicit-resource-declarations" task, ox-alpha passed on the first attempt, whereas GLM-5.3, GPT-5.6-sol, and Grok 4.6 had previously failed all four attempts (0/4). Meanwhile, ox-alpha maintained a 100% pass rate across 51,469 regression tests.

The report also documented an agent task involving 69 tool calls. Throughout the process, the model made only one error, did not enter any retry loops, and incurred low reasoning overhead. **Based on this, Davis believes that ox-alpha's performance is significantly stronger than GLM-5.3, resembling a next-generation model checkpoint rather than a simple variant.**

## Why Point to Z.AI?

In addition to technical fingerprints, Davis made cross-judgments based on the model's release timeline and Z.AI's previous testing methods.

**Z.AI released the text-only version of GLM-5.3 on August 14, while the unified vision flagship model has been a focus of community attention. More importantly, Z.AI has a precedent of testing models through stealth channels, with Pony Alpha eventually confirmed to be related to GLM-5.**

In terms of model scale, the report noted that ox-alpha's decoding speed differs from GLM-5V-Turbo by approximately 6%. Since the latter has 744 billion total parameters and 40 billion active parameters, Davis speculated that ox-alpha might adopt a Mixture of Experts (MoE) architecture of similar scale.

The report further suggested that if the model indeed has around 40 billion active parameters, the operator's claimed daily service capacity of 100 trillion tokens would be more feasible in terms of technology and cost.

Meanwhile, the report systematically ruled out other potential sources. Xiaomi's MiMo showed clear differences from ox-alpha in video encoder and audio interface; DeepSeek had not previously released video capabilities, and its tokenizer and model release method differed; Google, Qwen, xAI, OpenAI, and Anthropic were considered mismatched with ox-alpha in terms of tokenizer, output style, or video encoder.

## Free Testing May Continue Until August 27

It is worth noting that the current free access window for ox-alpha may last until August 27.

Davis pointed out that **some previous similar stealth models were officially claimed by relevant Chinese AI laboratories after their free testing periods ended.**

Currently, Z.AI has not issued an official response regarding the identity of ox-alpha. If the model is ultimately confirmed to be Z.AI's next-generation multimodal flagship, this "stealth test" on OpenRouter could serve as a public preview before its official launch.

Before official confirmation, the origin of ox-alpha remains inconclusive, but multiple independent tests currently point clues in the same direction, from the video encoder and tokenizer to model behavior.

### Related Stocks

- [02513.HK](https://longbridge.com/en/quote/02513.HK.md)

## Related News & Research

- [Z.AI Opens GLM-5.3 API Access; Model Weights to Be Open-Sourced Next Week](https://longbridge.com/en/news/296287311.md)
- [Alibaba’s lightweight Qwen takes on OpenAI, DeepSeek, Zhipu’s larger AI systems](https://longbridge.com/en/news/296228132.md)
- [17:00 ETCulture Amp Adds Missing Context to AI, Infusing Culture Intelligence Directly Inside the Tools Leaders Use Daily](https://longbridge.com/en/news/296828906.md)
- [Zhipu AI’s answer to Project Glasswing marks shift for Chinese cyber safety](https://longbridge.com/en/news/296208292.md)
- [CLSA: Z.AI  GLM-5.3 Post-training Capabilities Improve Significantly, Maintains Outperform Rating](https://longbridge.com/en/news/296314400.md)

---
> **Disclaimer: This article is for reference only and does not constitute any investment advice.**