AI Newsway

Mistral's 1T 'Le Chonk' Sells a Refusal Gap, Not a Leaderboard Win

Mistral Large 4 is in preview with 49 billion active parameters, and the pitch is security work that closed models decline to do

|5 min read0
AI Summary
Mistral AI opened a public preview of Mistral Large 4 on October 6, a 1-trillion-parameter mixture-of-experts model with 49 billion active parameters. Its headline claim is cybersecurity: 82% on a vulnerability reproduce-and-patch test where Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. Independent testing ranks the preview below US and Chinese flagships but above every Western open-weight rival. Weights are promised by the end of October.
An NVIDIA DGX system on display; Mistral trained Large 4 from scratch on roughly 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.
An NVIDIA DGX system on display; Mistral trained Large 4 from scratch on roughly 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.

Mistral AI put Mistral Large 4 into public preview on Tuesday, and the most interesting number in the launch post is not a benchmark score. It is a zero. The French lab says Claude Opus 5.5 and GPT-6 Astra both land near zero on a security test that asks a model to reproduce a real software vulnerability and then patch it, because they refuse the request. Mistral's model, nicknamed le Chonk, reports 82% on the same test.

Key takeaways

  • Mistral Large 4 is a 1-trillion-parameter mixture-of-experts model with 49 billion active parameters, trained from scratch on roughly 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.
  • The commercial argument is the refusal gap: Mistral claims a top-five global placing on the Artificial Analysis Cyber Index and 93% on Cybench's 40 exercises, in a category where safety-filtered closed models often decline to compete.
  • Weights are not out. Mistral promises them on Hugging Face by the end of October, after red-teaming, which leaves roughly three weeks where the "open" model is API-only.

Why a refusal is a product feature

Security teams have an awkward dependency problem. Proving a vulnerability is exploitable is the step that justifies patching it, and it is also the step that reads to a safety classifier like an attack. A model that declines mid-incident is not merely unhelpful; it removes a capability at the worst possible moment, which is the argument Mistral makes for pairing strong cyber performance with weights customers can host themselves.

That framing only works because the gap is measurable. Vendors rarely publish a rival's refusal rate, and Mistral is effectively using competitors' guardrails as a benchmark axis. It is a narrow axis, but a real one for regulated buyers, and it is the clearest reason to download a European model instead of calling an American endpoint.

Notably, Mistral also claims the opposite property at the same time. It reports the highest average refusal rate among open models on malicious cyber prompts drawn from JailbreakBench, StrongREJECT and AgentHarm, and 93.3% attack resistance on Lakera's public B3 security benchmark. Whether those two claims hold together under outside scrutiny is the thing to watch once weights ship.

How it scores everywhere else

Outside security the picture is competitive rather than dominant. Mistral's composite Coding Agent Index figure of 49.8%, built from 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4, edges past DeepSeek V4 Pro 0813 and Qwen3.8 Max by a margin small enough that a checkpoint refresh on either side would erase it. The agentic numbers are similar: 59.9% on AutomationBench's 657 business workflows, and 1,393 Elo on the AA-Briefcase knowledge-work set.

The most honest datapoint Mistral published is a blind human comparison run with Surge AI, where annotators scored anonymized coding outputs. Le Chonk came second of five at 3.74, behind Claude Opus 5 at 4.22, and only fractionally ahead of GLM-5.3 and Kimi K3. Human raters, in other words, put it in the same tier as the Chinese open models rather than above them.

Independent testing is blunter still. The Register reported that Artificial Analysis slots the preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna on aggregate intelligence, while noting it clearly beats Thinking Machines Lab's Inkling, the strongest American open-weight release. The win Mistral has actually earned is regional, not global.

The compute argument underneath

Mistral VP of science Pierre Stock told TechCrunch the run used about 4,000 NVIDIA GPUs, which he put at two to three times fewer than Chinese competitors and far below closed-source labs. Read alongside the launch post's figure of 3,800 Grace Blackwell chips, the claim is that a trillion-parameter model is now reachable from a mid-sized European cluster rather than a hyperscaler's.

Sparsity does most of that work. Activating 49 billion of a trillion parameters per token keeps serving costs closer to a mid-size dense model, and The Register judged the result small enough to run on a single eight-GPU server. Training data spanned more than 160 languages, including every official EU language — a detail aimed squarely at public-sector procurement.

What comes next

This is the first model funded by the €3 billion round Mistral closed last month at a €21 billion valuation, and the company says the reinforcement learning behind this checkpoint is still running with no sign of saturation. Expect a stronger revision before the benchmark comparisons above have time to settle.

FAQ

Is Mistral Large 4 open source?

Not today. Mistral has committed to publishing the weights by the end of October, which would make it an open-weight model rather than fully open source. The launch post set no license terms, and current access is limited to the preview API on Mistral Studio.

Why are the weights delayed?

Mistral says it is red-teaming the model first with cybersecurity leaders, vetted partners and state authorities, who test a build with reduced moderation. Stock said the company wants assurance the open weights are used for defense before release, which is an unusual gate for an open-weight launch and a direct consequence of leading on cyber capability.

How does it compare to GLM-5.3 and DeepSeek V4 Pro?

Close, by Mistral's own evidence. Its multimodal design wins the composite coding index and the cyber tests, ties on finance, and finishes within 0.15 points of GLM-5.3 in blind human grading. Artificial Analysis ranks the preview below the Chinese flagships on aggregate intelligence.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Silent Weight Swaps Are Testing What an API Model Name Guarantees
LLM & Chatbots

Silent Weight Swaps Are Testing What an API Model Name Guarantees

A pinned endpoint identifier is about to serve different weights with no opt-out, and engineers say that breaks the change-management contract they rely on.

Seung Jung27 days ago
IBM Bets Against the Architecture Crowd With Granite 4.2 Reasoning Models
LLM & Chatbots

IBM Bets Against the Architecture Crowd With Granite 4.2 Reasoning Models

IBM's Granite 4.2 ships 3B, 8B and 30B dense reasoning models under Apache 2.0 with a 512K context window and an agentic RL stage for the larger two.

Seung Jung42 days ago
Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing
LLM & Chatbots

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

Z.ai confirmed the stealth Ox Alpha model is GLM-5.3-Flash: 320B parameters, 18B active, MIT-licensed weights and sub-dollar output pricing.

Seung Jung40 days ago
Researchers Bypass Grok's Guardrails by Encrypting the Attack Payload
LLM & Chatbots

Researchers Bypass Grok's Guardrails by Encrypting the Attack Payload

Adversa researchers bypassed Grok's safety filters by encrypting malicious instructions with AES-256-GCM, letting the model decrypt and execute them itself.

Seung Jung47 days ago
GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use
LLM & Chatbots

GPT-6 Astra Goes to Work: OpenAI's Priciest Model Bets Everything on Computer Use

OpenAI has begun rolling out GPT-6 Astra to business customers, framing its newest frontier model less as a chatbot and more as a worker that operates software...

Seung Jung27 days ago
ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas
LLM & Chatbots

ChatGPT Images 2.5 Halves Generation Latency and Adds a Sketch Canvas

OpenAI shipped ChatGPT Images 2.5 with up to 50% lower latency, a Sketch drawing canvas, templates, pinned comments and two new API models.

Seung Jung26 days ago