AI Newsway

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

The anonymous model that topped OpenRouter was a preview, and Z.ai says it ran entirely on Chinese AI chips

|4 min read0
AI Summary
Z.ai confirmed on August 26 that the anonymous stealth/ox-alpha model that topped OpenRouter for six days is GLM-5.3-Flash, released the same night on Hugging Face under an MIT license. The sparse mixture-of-experts model carries 320 billion parameters with about 18 billion active, targets a one-million-token context, and lists at 15 cents per million input tokens and 50 cents per million output tokens. Vendor benchmarks show big gains over GLM-5.2, but GPT-5.6 Terra still leads on DeepSWE and Terminal Bench.
Server racks in a data center, the kind of inference capacity Z.ai says served its GLM-5.3-Flash preview entirely on Chinese AI chips
Server racks in a data center, the kind of inference capacity Z.ai says served its GLM-5.3-Flash preview entirely on Chinese AI chips

For six days in late August, the most talked-about model on OpenRouter had no name and no owner. It appeared as stealth/ox-alpha, offered a one-million-token context window, cost nothing to use, and climbed straight to the top of the leaderboard. Developers ran forensics on it. Stack traces, an obscure error code, and tokenizer comparisons all pointed one direction.

They were right. Ox Alpha belonged to Z.ai, the Chinese lab behind the GLM series. Bloomberg reported the connection on the morning of August 26. That evening, Z.ai confirmed it and gave the model its production name: GLM-5.3-Flash.

Weights went up on Hugging Face the same night under an MIT license. API documentation went live alongside them. The stealth listing had been a preview skin for a product that was already finished.

The specifications behind the leaderboard run

GLM-5.3-Flash is a sparse mixture-of-experts design. It carries 320 billion total parameters but activates roughly 18 billion per token. That ratio is the entire economic argument for the model.

It is natively multimodal across text, image, and video rather than a text model with vision attached afterward. Z.ai frames vision as part of the coding loop, where a model renders output, checks it, and refines. The context target is one million tokens, supported by a hybrid linear and sparse attention scheme.

The company claims that architecture cuts serving cost substantially against its own larger GLM-5.3, citing roughly three times less attention compute and a KV cache about 4.4 times smaller on per-token metrics. Pre-training ran on a multimodal corpus the lab describes as 30 trillion tokens.

Pricing is where the pressure lands. List rates are 15 cents per million input tokens and 50 cents per million output tokens, with cached input at 3 cents. Those are Flash-tier numbers for a model being positioned against frontier systems.

Reading the benchmark table carefully

Z.ai published an evaluation table at launch, and it deserves a skeptical eye because the vendor produced it. The results are also not a clean sweep.

The clearest gains are against the lab's own previous generation. On DeepSWE the model reports 63.4 against 46.2 for GLM-5.2. On AutomationBench it reports 48.8 against 26.2. Both are agentic workloads, which matches the traffic Ox Alpha was absorbing on OpenRouter.

Against outside competition the picture is mixed. Z.ai reports leading GDPVal-AA v2, placing it ahead of Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash on that measure. But GPT-5.6 Terra still leads on DeepSWE and Terminal Bench in the same table, and Gemini 3.7 Flash leads AutomationBench.

One widely shared figure from the stealth period, an 80 percent DeepSWE score, sits far above the official 63.4. The gap most likely reflects a different evaluation harness, which is a useful reminder that agent benchmarks are not portable between setups.

The line that will get quoted

Z.ai added one detail to its announcement that carries more weight than any benchmark row. The entire stealth preview, it said, was served on Chinese AI chips.

That claim reframes the release. It is no longer only a story about a cheap capable model. It is a claim about domestic inference capacity at leaderboard-topping scale, made at a moment when export controls are meant to constrain exactly that.

The commercial threat to expensive frontier vendors is straightforward. Open weights under MIT terms let teams self-host and fine-tune without a licensing conversation. Launch-day support for SGLang, vLLM, and TokenSpeed removes most of the integration friction. A sub-dollar output price removes the rest.

The playbook itself is worth noting, because Z.ai has now demonstrated it works. Ship anonymously, let a developer community validate the model without brand bias, become the most used model of the week, then attach the name and release the weights. Earned credibility arrives before the marketing does.

Anyone building on the preview should migrate off the stealth identifier and route to the named model or a self-hosted copy, since preview endpoints are not guaranteed to persist.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung29 days ago
Grok 4.6 Reaches the AI Frontier Without Raising Its Price
LLM & Chatbots

Grok 4.6 Reaches the AI Frontier Without Raising Its Price

xAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol while holding pricing flat at $2/$6 per million tokens.

Seung Jung34 days ago
Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two
LLM & Chatbots

Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two

Alibaba's Qwen3.8 Max leads the Artificial Analysis agentic index, winning through long-horizon persistence rather than top reasoning scores.

Seung Jung41 days ago
Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price
LLM & Chatbots

Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price

Google's Gemini 3.7 Flash lands three weeks after 3.6 Flash, scoring 65.3% on DeepSWE and shipping at half the price through the end of 2026.

Seung Jung34 days ago
DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing
LLM & Chatbots

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Seung Jung34 days ago
Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse
AI & Machine Learning

Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse

Alibaba's Qwen team released Qwen3.8 as open weights, pairing a 2.4-trillion-parameter MoE flagship with a compact 27B vision-language model.

Seung Jung33 days ago