Mistral AI put Mistral Large 4 into public preview on Tuesday, and the most interesting number in the launch post is not a benchmark score. It is a zero. The French lab says Claude Opus 5.5 and GPT-6 Astra both land near zero on a security test that asks a model to reproduce a real software vulnerability and then patch it, because they refuse the request. Mistral's model, nicknamed le Chonk, reports 82% on the same test.
Key takeaways
- Mistral Large 4 is a 1-trillion-parameter mixture-of-experts model with 49 billion active parameters, trained from scratch on roughly 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.
- The commercial argument is the refusal gap: Mistral claims a top-five global placing on the Artificial Analysis Cyber Index and 93% on Cybench's 40 exercises, in a category where safety-filtered closed models often decline to compete.
- Weights are not out. Mistral promises them on Hugging Face by the end of October, after red-teaming, which leaves roughly three weeks where the "open" model is API-only.
Why a refusal is a product feature
Security teams have an awkward dependency problem. Proving a vulnerability is exploitable is the step that justifies patching it, and it is also the step that reads to a safety classifier like an attack. A model that declines mid-incident is not merely unhelpful; it removes a capability at the worst possible moment, which is the argument Mistral makes for pairing strong cyber performance with weights customers can host themselves.
That framing only works because the gap is measurable. Vendors rarely publish a rival's refusal rate, and Mistral is effectively using competitors' guardrails as a benchmark axis. It is a narrow axis, but a real one for regulated buyers, and it is the clearest reason to download a European model instead of calling an American endpoint.
Notably, Mistral also claims the opposite property at the same time. It reports the highest average refusal rate among open models on malicious cyber prompts drawn from JailbreakBench, StrongREJECT and AgentHarm, and 93.3% attack resistance on Lakera's public B3 security benchmark. Whether those two claims hold together under outside scrutiny is the thing to watch once weights ship.
How it scores everywhere else
Outside security the picture is competitive rather than dominant. Mistral's composite Coding Agent Index figure of 49.8%, built from 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4, edges past DeepSeek V4 Pro 0813 and Qwen3.8 Max by a margin small enough that a checkpoint refresh on either side would erase it. The agentic numbers are similar: 59.9% on AutomationBench's 657 business workflows, and 1,393 Elo on the AA-Briefcase knowledge-work set.
The most honest datapoint Mistral published is a blind human comparison run with Surge AI, where annotators scored anonymized coding outputs. Le Chonk came second of five at 3.74, behind Claude Opus 5 at 4.22, and only fractionally ahead of GLM-5.3 and Kimi K3. Human raters, in other words, put it in the same tier as the Chinese open models rather than above them.
Independent testing is blunter still. The Register reported that Artificial Analysis slots the preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna on aggregate intelligence, while noting it clearly beats Thinking Machines Lab's Inkling, the strongest American open-weight release. The win Mistral has actually earned is regional, not global.
The compute argument underneath
Mistral VP of science Pierre Stock told TechCrunch the run used about 4,000 NVIDIA GPUs, which he put at two to three times fewer than Chinese competitors and far below closed-source labs. Read alongside the launch post's figure of 3,800 Grace Blackwell chips, the claim is that a trillion-parameter model is now reachable from a mid-sized European cluster rather than a hyperscaler's.
Sparsity does most of that work. Activating 49 billion of a trillion parameters per token keeps serving costs closer to a mid-size dense model, and The Register judged the result small enough to run on a single eight-GPU server. Training data spanned more than 160 languages, including every official EU language — a detail aimed squarely at public-sector procurement.
What comes next
This is the first model funded by the €3 billion round Mistral closed last month at a €21 billion valuation, and the company says the reinforcement learning behind this checkpoint is still running with no sign of saturation. Expect a stronger revision before the benchmark comparisons above have time to settle.
FAQ
Is Mistral Large 4 open source?
Not today. Mistral has committed to publishing the weights by the end of October, which would make it an open-weight model rather than fully open source. The launch post set no license terms, and current access is limited to the preview API on Mistral Studio.
Why are the weights delayed?
Mistral says it is red-teaming the model first with cybersecurity leaders, vetted partners and state authorities, who test a build with reduced moderation. Stock said the company wants assurance the open weights are used for defense before release, which is an unusual gate for an open-weight launch and a direct consequence of leading on cyber capability.
How does it compare to GLM-5.3 and DeepSeek V4 Pro?
Close, by Mistral's own evidence. Its multimodal design wins the composite coding index and the cyber tests, ties on finance, and finishes within 0.15 points of GLM-5.3 in blind human grading. Artificial Analysis ranks the preview below the Chinese flagships on aggregate intelligence.






