AI Newsway

Sonnet 5.5 Sends Blocked Cyber Requests Back to Sonnet 5 β€” and API Users Have to Opt In

Anthropic's own testing puts the injection compromise rate on rerouted traffic roughly 175 times higher than on requests Sonnet 5.5 answers itself

|5 min read0
AI Summary
Claude Sonnet 5.5 is the first Sonnet model to ship under Anthropic's Opus-tier cyber policy, which can reroute blocked cybersecurity requests to Sonnet 5. Claude's apps retry automatically while API developers must opt in, and without fallback a blocked request fails. Anthropic's prompt-injection testing compromised 12.01% of rerouted requests, against four of 5,901 answered by Sonnet 5.5 itself β€” a reason to treat fallback as a security setting.
A user interacting with a chat assistant, the interface layer where Anthropic's Claude apps display a notice when a blocked request falls back to Sonnet 5
A user interacting with a chat assistant, the interface layer where Anthropic's Claude apps display a notice when a blocked request falls back to Sonnet 5

The riskiest line in Claude Sonnet 5.5's launch material is not a capability score. It is a routing rule: when a classifier blocks a cybersecurity request, the platform may answer it with the older Sonnet 5 instead, and Anthropic's own prompt-injection testing found rerouted traffic far easier to compromise. The New Stack first laid out how the behaviour differs between Claude's apps and the API.

Key takeaways

  • Claude's apps reroute blocked cyber requests to Sonnet 5 automatically, while API callers must enable fallback themselves β€” and with it off, a blocked request fails instead of being answered.
  • Anthropic's coding-environment tests compromised 12.01% of rerouted requests, against four of the 5,901 requests Sonnet 5.5 answered itself, a gap of roughly two orders of magnitude.
  • Sonnet 5.5 is the first model in the Sonnet tier to launch under the cyber policy Anthropic previously applied only to Opus.

What changes in the request path

Two things are new, and only one is visible in a model card. The first is the block itself: a narrow set of cyber requests, plus a slim category tied to frontier-model development such as kernel work on particular ML accelerators, can be refused outright.

The second is what happens next. Inside Anthropic's apps, a refused cyber request is retried on Sonnet 5 and the user sees a notice naming the model that answered. On the API that retry does not exist until a developer enables it. For anyone treating the release as a version bump in a config file, that is the trap: a pipeline that worked on Sonnet 5 can start returning hard failures for the same security-adjacent prompts.

Why fallback is a security decision, not a reliability one

Enabling fallback looks like a straightforward availability win. Anthropic's numbers say otherwise. In prompt-injection testing of coding environments, a quarter of requests aimed at Sonnet 5.5 ended up served by Sonnet 5 after tripping the cyber classifier β€” often because injected instructions to wipe disks or delete files resemble precisely the behaviour the classifier exists to stop.

Of that rerouted slice, 12.01% were successfully compromised. Sonnet 5.5, handling requests itself, was compromised four times across 5,901 attempts, or about 0.07%. Stated plainly: turning fallback on can route a quarter of adversarial traffic to a model that failed injection tests at more than a hundred times the rate of the one you selected. A separate indirect-injection benchmark from guardrail vendor Gray Swan showed no drop with fallback enabled, so the effect is not uniform across suites.

The capability jump that triggered the policy

None of this was applied speculatively. On Anthropic's published figures, Sonnet 5.5 scores 70.6% on the Terminal-Bench 4.0 agentic coding test, against 10.3% for Sonnet 5 and 66.4% for the costlier Opus 5.5, at unchanged pricing of $2 per million input tokens and $10 per million output.

The offensive-security gains were sharper. With safeguards switched off, the system card records full arbitrary code execution in 178 of 410 ExploitBench runs. On Irregular's CyScenarioBench the model completed 46.1% of challenges where Sonnet 5 managed 0.7%, and on a binary exploitation benchmark derived from Google's OSS-Fuzz corpus it produced 50 control-flow hijacks against three. Anthropic still ranks it below Opus 5.5 on cyber capability; the gap from the previous Sonnet was simply wide enough to pull the Opus-tier policy down a price tier.

How the block is enforced

Three components decide. A probe inspects the model's internal activations, a lightweight classifier runs on Sonnet 5.5 itself, and a separately trained LLM classifier weighs the probe's verdict before a conversation is cut. Anthropic describes cyber detection as comparable to Opus 5's with deliberately softer jailbreak protection, and tells users to expect more refusals than Sonnet 5 produced β€” legitimate security work included. A company representative cited penetration testing, exploit generation and binary vulnerability scanning as in-scope examples.

What to settle before the upgrade

  1. Decide the fallback setting deliberately, and document it as a security choice rather than a retry policy.
  2. Log which model answered each request; Anthropic says rerouted responses identify themselves.
  3. Re-run injection testing against Sonnet 5, not only the model you nominally selected.

Classifier tuning to cut false positives continues, and verified defenders are to get looser access through an expanded Cyber Verification Program that Sonnet 5.5 had not joined at launch. Our earlier report covered the benchmark side of the same release.

FAQ

Is the Sonnet 5 fallback enabled by default?

Not on the API. Anthropic's Claude apps retry blocked cyber requests on Sonnet 5 automatically, but API developers have to switch the fallback on, and third-party platforms may behave differently again. With it off, a blocked request fails rather than being rerouted.

Will ordinary coding work hit these safeguards?

Anthropic says most requests will not and that routine software development is unaffected. The policy is aimed at higher-risk cybersecurity tasks such as penetration testing and exploit generation, though the company also warns of more refusals than Sonnet 5 on legitimate security work.

Does Sonnet 5.5 cost more than Sonnet 5?

No. Pricing is identical at $2 per million input tokens and $10 per million output tokens, and Anthropic says the model typically needs fewer tokens per task, putting effective cost up to 30% below its predecessor.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Claude Code Projects Returns as a Coordinator That Runs Parallel Cloud Threads
Developer Tools

Claude Code Projects Returns as a Coordinator That Runs Parallel Cloud Threads

Anthropic's redesigned Claude Code Projects puts a coordinator above worker threads, each a full cloud session on its own branch, with shared memory.

Seung Jung11 days ago
A Million Agent Skills in Seven Months β€” and Nearly Half Were Installed Exactly Once
Developer Tools

A Million Agent Skills in Seven Months β€” and Nearly Half Were Installed Exactly Once

Vercel's skills.sh registry hit 1M agent skills in seven months. Nearly half were installed once, while 375 listings account for 62% of all installs.

Seung Jung3 days ago
Researchers Reached OpenAI's Internal Repo Through a Forum Image Bug
Developer Tools

Researchers Reached OpenAI's Internal Repo Through a Forum Image Bug

Hacktron AI chained a libheif heap overflow with an OpenAI SSO flaw to reach employee Codex accounts and the openai/openai monorepo. Both bugs are patched.

Seung Jung11 days ago
3,000 Merged Changes, Zero Rollbacks: Inside Anthropic's Two-Week Speed Sprint
Developer Tools

3,000 Merged Changes, Zero Rollbacks: Inside Anthropic's Two-Week Speed Sprint

Anthropic's August sprint cut claude.ai's p75 time-to-typeable from 3.1s to 0.55s, merging 3,000 changes with no rollback by gating CI on instruction counts.

Seung Jung5 days ago
Bun's Zig-to-Rust Port Took 11 Days and 64 Parallel Agents
Developer Tools

Bun's Zig-to-Rust Port Took 11 Days and 64 Parallel Agents

Bun's runtime moved from Zig to Rust in an 11-day agent-run port. The memory-safety bugs vanished; the meaning bugs did not.

Seung Jung8 days ago
Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5
Developer Tools

Four Opus 5 Request Patterns Now Return 400 on Claude Opus 5.5

Claude Opus 5.5 is 20% cheaper than Opus 5 and returns HTTP 400 on four request patterns Opus 5 accepted. A fifth change silences agent progress streams.

Seung Jungyesterday