AI Newsway

Qwen3.8 Max Tops Artificial Analysis Agentic Index, Outranking Every US Lab but Two

Alibaba's flagship model wins on persistence rather than raw reasoning, averaging 64 turns per task where its predecessor took 14

|3 min read0
AI Summary
Alibaba's Qwen3.8 Max took the top spot on Artificial Analysis's agentic index, a benchmark measuring multi-step task execution, outranking every model except top releases from Anthropic and OpenAI despite a mid-pack score of 56 on the firm's Intelligence Index. The model averages 64 turns to finish GDPval-AA tasks versus 14 for its predecessor, showing persistence, not reasoning, drives its lead, and it beats GPT-5.6 Sol Max on the OSWorld-Verified benchmark. A Chinese model now tops a leaderboard US labs controlled.
The Qwen logo, marking Alibaba's large language model family whose Qwen3.8 Max flagship now leads the Artificial Analysis agentic index
The Qwen logo, marking Alibaba's large language model family whose Qwen3.8 Max flagship now leads the Artificial Analysis agentic index

Alibaba's Qwen3.8 Max has taken the number one position on the agentic index maintained by Artificial Analysis, an independent evaluation firm that scores how reliably a model plans and executes multi-step work rather than how well it answers a single question. The result puts a Chinese-developed system at the top of a leaderboard that American labs have largely controlled for the past two years.

The ranking spread quickly through developer circles, collecting close to 300 points and more than 180 comments on Hacker News within hours of surfacing. Much of that discussion centered less on the headline placement than on the unusual way the model appears to have earned it.

Persistence, Not Raw Intelligence

On the broader Artificial Analysis Intelligence Index, which weighs reasoning and knowledge instead of agentic execution, Qwen3.8 Max scores 56. That figure sits well above the median of 32 for reasoning models in a comparable price bracket and draws level with Anthropic's Claude Opus 4.8. It also clears every model shipped by Google, Meta and xAI. Only the top-tier releases from Anthropic and OpenAI, plus Moonshot AI's Kimi K3, currently rank higher.

The distance between a mid-pack intelligence score and a first-place agentic finish points to something other than raw capability. Evaluation data shows Qwen3.8 Max averaging 64 turns to finish tasks on the GDPval-AA benchmark, against just 14 turns for its predecessor, Qwen3.7 Max. The newer model is markedly more willing to keep grinding at a problem instead of declaring an early answer.

That behavior is a deliberate design choice with real trade-offs. Long-horizon persistence is precisely what agentic workloads reward, since a coding or research agent that abandons a task halfway is worth little regardless of how sharp its individual reasoning steps are. It also means substantially more tokens consumed per task, which shifts where the model makes economic sense.

Strong Showing on Computer Use

The agentic gains show up elsewhere too. Qwen3.8 Max posts 86.1 on OSWorld-Verified, a benchmark that measures how capably an agent can drive an operating system and its applications. That places it ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, a narrow but meaningful lead in a category that has proven stubbornly difficult for frontier models.

Computer use has become a proxy for practical agent readiness. Navigating menus, handling unexpected dialogs and recovering from misclicks demand exactly the tolerance for long, messy sequences that the turn-count data suggests Qwen was tuned for.

What It Signals

Alibaba released Qwen3.8 Max alongside smaller open-weight models aimed at coding and collaborative work, continuing a strategy of pairing a frontier proprietary flagship with freely available siblings. That combination has steadily widened Qwen's footprint among developers who want capable models they can self-host.

Benchmark leadership is rarely durable, and Artificial Analysis rankings shift with each major release. The more durable signal is what the underlying numbers reveal about how the gap is closing. Qwen3.8 Max did not win by matching the strongest US models on reasoning; it won by being tuned harder for the specific shape of agentic work. That is a repeatable strategy, and one competitors will now have to answer.

For teams evaluating models for agent deployments, the takeaway is that headline intelligence scores are an increasingly poor proxy for agentic reliability. The two are drifting apart, and the benchmarks are only beginning to reflect it.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse
AI & Machine Learning

Qwen3.8 Goes Open: Alibaba Ships a 2.4T Flagship and a 27B Workhorse

Alibaba's Qwen team released Qwen3.8 as open weights, pairing a 2.4-trillion-parameter MoE flagship with a compact 27B vision-language model.

Seung Jung33 days ago
GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price
LLM & Chatbots

GLM-5.3 Scores 60 on Artificial Analysis Index at Half the Usual Output Price

Artificial Analysis scored Z.ai's GLM-5.3 at 60 on its Intelligence Index, well above the 35 median, at $4.40 per million output tokens. The catch is verbosity.

Seung Jung29 days ago
Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price
LLM & Chatbots

Gemini 3.7 Flash Arrives Three Weeks After 3.6 - At Half the Price

Google's Gemini 3.7 Flash lands three weeks after 3.6 Flash, scoring 65.3% on DeepSWE and shipping at half the price through the end of 2026.

Seung Jung34 days ago
OpenAI Drops Message Caps on ChatGPT's Free Tier and Hands It a Think Button
LLM & Chatbots

OpenAI Drops Message Caps on ChatGPT's Free Tier and Hands It a Think Button

OpenAI is lifting message caps on ChatGPT's free tier and adding a Think button, with GPT-5.6 Luna becoming the default for Free and Go users.

Seung Jung41 days ago
Qwen3.8's 27B Open Model Is the Release That Actually Matters
LLM & Chatbots

Qwen3.8's 27B Open Model Is the Release That Actually Matters

Alibaba's Qwen3.8 open weights pair a 2.4T mixture-of-experts model with a 27B dense multimodal model sized for a single 24GB consumer GPU.

Seung Jung33 days ago
DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing
LLM & Chatbots

DeepSeek V4 Pro Hits General Availability at a Fraction of Frontier Pricing

DeepSeek has moved its V4 Pro model to general availability, and the release is drawing attention less for raw capability than for what that capability now cost...

Seung Jung34 days ago