Independent evaluator Artificial Analysis has published measurements for GLM-5.3 (max), the reasoning model Z.ai released on August 14, and the results place the Chinese lab firmly inside frontier territory. The model scores 60 on the Artificial Analysis Intelligence Index against a median of 35 for comparable systems, while charging $1.40 per million input tokens and $4.40 per million output tokens.
Both prices land under the field. The median input price among tracked models in the same class is $1.75 and the median output price is $10.00, so GLM-5.3 undercuts typical output billing by more than half while scoring well above the middle of the pack. Throughput is also above average at 85 tokens per second versus a 74-token norm, and the context window runs to one million tokens.
Verbosity is where the discount leaks
The complication appears in the token accounting. Completing the index consumed 170 million output tokens from GLM-5.3, more than double the 72 million median, and the full evaluation ran to $1,238.50. Cheap per token is not the same as cheap per finished task: a model that reasons at twice the typical length can spend its price advantage before it returns an answer, which matters most for agent loops billed on completion rather than on request count.
Version 4.1.1 of the index aggregates nine evaluations spanning GDPval-AA v2, tau-3-Banking, Terminal-Bench v2.1, SciCode, GPQA Diamond, Humanity's Last Exam, CritPt, AA-LCR and AA-Omniscience. That last test is the interesting one for production buyers, because it rewards correct answers, penalizes confident errors and applies no penalty at all for declining to answer, scoring reliability rather than raw recall.
Post-training, not a new base model
How Z.ai reached the score is arguably more notable than the score. GLM-5.3 sits on the same base model as GLM-5.2, with the lab stating flatly that scaling post-training was the entire recipe: more environments, more diverse tasks and more compute spent on reinforcement learning rather than another pretraining run. At roughly 750 billion parameters the model is about a third the size of Moonshot AI's Kimi K3, which it passes on several benchmarks.
Z.ai's own figures claim a 50% gain over GLM-5.2 on its in-house code benchmark and a leading 84.5% on the CyberGym defensive-security evaluation, ahead of GPT-5.6 Sol at 83.6%. Those are self-reported and have not been reproduced by a third party, which is a meaningful gap given that the weights are not out. GLM-5.3 launched inside Z.ai's coding subscription first, with API access to follow and a Hugging Face release expected roughly two weeks after launch, held pending a cybersecurity review. Artificial Analysis accordingly lists the max configuration as proprietary rather than open-weights for now.
That delay is its own signal. A lab confident enough to publish frontier coding and security numbers while withholding weights for safety screening is behaving like an incumbent, not a fast follower, and it deprives the open-source community of the usual verification path in the window where the claims matter most.
For teams evaluating models, the practical read is narrower than the headline. An index score compresses nine very different tasks into one number, and no aggregate predicts how a model behaves inside a specific toolchain, prompt format or latency budget. The durable finding is structural: a mid-sized model reached this tier through post-training investment alone, at prices that put continued pressure on Western labs charging premium rates for reasoning tokens. Whether that pressure translates into repricing depends on how many buyers are willing to run a model whose weights arrive on the vendor's schedule.






