AI Newsway

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Independent testing puts Z.ai's newest reasoning model near the frontier, but its verbosity complicates the cheap-tokens story

|3 min read0
AI Summary
Artificial Analysis measured Z.ai's GLM-5.3, released August 14, at 60 on its Intelligence Index against a median of 35, priced at $1.40 per million input and $4.40 per million output tokens versus medians of $1.75 and $10.00. The catch is verbosity: the model consumed 170 million output tokens to complete the index, more than double the 72 million median, costing $1,238.50. Watch API availability and whether third parties reproduce Z.ai's self-reported CyberGym lead.
Source code on a developer display, the workload class Z.ai optimized GLM-5.3 for through extended post-training.
Source code on a developer display, the workload class Z.ai optimized GLM-5.3 for through extended post-training.

Independent evaluator Artificial Analysis has published measurements for GLM-5.3 (max), the reasoning model Z.ai released on August 14, and the results place the Chinese lab firmly inside frontier territory. The model scores 60 on the Artificial Analysis Intelligence Index against a median of 35 for comparable systems, while charging $1.40 per million input tokens and $4.40 per million output tokens.

Both prices land under the field. The median input price among tracked models in the same class is $1.75 and the median output price is $10.00, so GLM-5.3 undercuts typical output billing by more than half while scoring well above the middle of the pack. Throughput is also above average at 85 tokens per second versus a 74-token norm, and the context window runs to one million tokens.

Verbosity is where the discount leaks

The complication appears in the token accounting. Completing the index consumed 170 million output tokens from GLM-5.3, more than double the 72 million median, and the full evaluation ran to $1,238.50. Cheap per token is not the same as cheap per finished task: a model that reasons at twice the typical length can spend its price advantage before it returns an answer, which matters most for agent loops billed on completion rather than on request count.

Version 4.1.1 of the index aggregates nine evaluations spanning GDPval-AA v2, tau-3-Banking, Terminal-Bench v2.1, SciCode, GPQA Diamond, Humanity's Last Exam, CritPt, AA-LCR and AA-Omniscience. That last test is the interesting one for production buyers, because it rewards correct answers, penalizes confident errors and applies no penalty at all for declining to answer, scoring reliability rather than raw recall.

Post-training, not a new base model

How Z.ai reached the score is arguably more notable than the score. GLM-5.3 sits on the same base model as GLM-5.2, with the lab stating flatly that scaling post-training was the entire recipe: more environments, more diverse tasks and more compute spent on reinforcement learning rather than another pretraining run. At roughly 750 billion parameters the model is about a third the size of Moonshot AI's Kimi K3, which it passes on several benchmarks.

Z.ai's own figures claim a 50% gain over GLM-5.2 on its in-house code benchmark and a leading 84.5% on the CyberGym defensive-security evaluation, ahead of GPT-5.6 Sol at 83.6%. Those are self-reported and have not been reproduced by a third party, which is a meaningful gap given that the weights are not out. GLM-5.3 launched inside Z.ai's coding subscription first, with API access to follow and a Hugging Face release expected roughly two weeks after launch, held pending a cybersecurity review. Artificial Analysis accordingly lists the max configuration as proprietary rather than open-weights for now.

That delay is its own signal. A lab confident enough to publish frontier coding and security numbers while withholding weights for safety screening is behaving like an incumbent, not a fast follower, and it deprives the open-source community of the usual verification path in the window where the claims matter most.

For teams evaluating models, the practical read is narrower than the headline. An index score compresses nine very different tasks into one number, and no aggregate predicts how a model behaves inside a specific toolchain, prompt format or latency budget. The durable finding is structural: a mid-sized model reached this tier through post-training investment alone, at prices that put continued pressure on Western labs charging premium rates for reasoning tokens. Whether that pressure translates into repricing depends on how many buyers are willing to run a model whose weights arrive on the vendor's schedule.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing
LLM & Chatbots

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Seung Jung34 days ago
Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing
LLM & Chatbots

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

Z.ai confirmed the stealth Ox Alpha model is GLM-5.3-Flash: 320B parameters, 18B active, MIT-licensed weights and sub-dollar output pricing.

Seung Jung20 days ago
Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse
AI & Machine Learning

Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse

Alibaba's Qwen team released Qwen3.8 as open weights, pairing a 2.4-trillion-parameter MoE flagship with a compact 27B vision-language model.

Seung Jung33 days ago
Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two
LLM & Chatbots

Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two

Alibaba's Qwen3.8 Max leads the Artificial Analysis agentic index, winning through long-horizon persistence rather than top reasoning scores.

Seung Jung41 days ago
Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price
LLM & Chatbots

Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price

Google's Gemini 3.7 Flash lands three weeks after 3.6 Flash, scoring 65.3% on DeepSWE and shipping at half the price through the end of 2026.

Seung Jung34 days ago
Silent Weight Swaps Are Testing What an API Model Name Guarantees
LLM & Chatbots

Silent Weight Swaps Are Testing What an API Model Name Guarantees

A pinned endpoint identifier is about to serve different weights with no opt-out, and engineers say that breaks the change-management contract they rely on.

Seung Jung7 days ago