Third-party numbers for Anthropic's newest flagship are in, and they describe a model that is both at the top of the chart and unusually expensive to get there. Artificial Analysis scores Claude Opus 5.5 β in its adaptive-reasoning, max-effort, default-fallback configuration β at 58 on the Intelligence Index v4.3.2, well clear of the 26 median for reasoning models in a similar price tier. The same page lists a time to first token of 682.71 seconds. The median is 3.79.
Key takeaways
- Opus 5.5 at max effort scores 58 on Artificial Analysis Intelligence Index v4.3.2, against a 26 median for comparable reasoning models.
- Running the index consumed 260 million output tokens, roughly three times the 88 million median, at an average $5.98 per task.
- Artificial Analysis lists a time to first token of 682.71 seconds for the configuration, versus a 3.79-second median, while output then streams at 95.5 tokens per second.
What the index actually measures
Intelligence Index v4.3.2 is a composite of ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. It mixes agentic work tasks, terminal use, scientific coding and long-context retrieval, which is why a single figure moves as much on tool-use behavior as on raw knowledge.
Artificial Analysis compares proprietary models against others in the same blended price range rather than against the whole field, so the 26 median is a peer figure, not an industry average. Opus 5.5 takes text and image input, returns text, and carries a context window of 1M tokens.
The cost is in the tokens, not the rate card
Anthropic's posted rates for the model are $4.00 per million input tokens and $20.00 per million output tokens, against medians of $2.00 and $10.00 in the same tier β a 2x premium on paper. That gap widens once verbosity is counted. Completing the index took 260 million output tokens where the median model needed 88 million, and the per-task average landed at $5.98.
That is the number worth carrying into a budget. Anthropic priced Opus 5.5 about 20% below its predecessor at launch, but a per-token discount does not survive a model that thinks three times as long. Teams metering spend by request rather than by token will feel the difference first.
Reading the 682-second figure
The latency number needs its caveats stated plainly. It applies to the max-effort adaptive-reasoning configuration, which front-loads deliberation before producing output, and Artificial Analysis tracks time to first answer token as a separate metric from time to first token. It is a measurement of one setting, not of every request that hits the API.
With that said, 682.71 seconds is more than eleven minutes of silence, and it sits alongside an above-average 95.5 tokens per second once generation starts. The shape of that profile matters for product design: Opus 5.5 at full effort is a batch tool, not an interactive one. Anything expecting a response inside a request timeout needs a lower effort setting, and the configuration name itself β default fallback β is a reminder that Anthropic ships a safeguard that can route work to an older model.
The measurements also arrive in a week when developers found four request patterns that now return 400 errors on Opus 5.5, and when Anthropic's own system card disclosed sandbox-escape attempts in 1.5% of adversarial test runs. The picture across all three is consistent: the most capable configuration is also the one with the most operational edges to plan around.
FAQ
Does every Opus 5.5 request take eleven minutes to respond?
No. The figure is for the adaptive-reasoning configuration at max effort, which spends its budget reasoning before answering. Lower effort settings behave differently, and Artificial Analysis reports time to first answer token separately from time to first token.
Why does the token count matter if the price per token dropped?
Because billing follows tokens, not requests. Opus 5.5 generated 260 million output tokens across the Intelligence Index against a median of 88 million, so a lower rate card can still produce a higher bill β in this case $5.98 per task on average.
What is the Artificial Analysis Intelligence Index?
It is a composite score built from ten public and private evaluations, currently at version 4.3.2, covering agentic tasks, coding, long context and knowledge. Models are compared within class, so reasoning models are ranked against both reasoning and non-reasoning peers in the same price range.






