← All posts

Artificial intelligence

Mistral's Large 4 Is the Best Model in the World at Reproducing Software Vulnerabilities. The Weights Go Public at the End of the Month

Mistral has previewed Large 4, a 1.05-trillion-parameter open-weight model, and the capability it leads on is security work that Claude Opus 5.5 and GPT-6 Astra refuse to do. Mistral scores 82 per cent on a vulnerability-reproduction test where those two score near zero — and says it will publish the weights within three weeks.

MAI
The cover image Mistral published with its own announcement of Mistral Large 4 on mistral.ai.

Mistral previewed Mistral Large 4 on Tuesday: 1.05 trillion total parameters, 49 billion active per token, a granular mixture-of-experts model with a one-million-token context window and a 1.6-billion-parameter vision encoder. It is in research public preview through Mistral's API as mistral-large-4-0, at $1.36 per million input tokens and $4.18 per million output. The company says the weights follow by the end of October.

The parameter count is the headline everyone reached for. It is not the interesting part. The interesting part is which capability Mistral chose to lead with, and what it means that the weights are going out.

The gap Mistral is selling is a policy gap, not a capability gap

On the Artificial Analysis Cyber Index, an independent evaluation of how well models find and fix flaws in real software, Mistral says Large 4 ranks in the global top five and leads open-weight models developed outside China by a wide margin. On one component — reproducing a vulnerability to show it is real — it scores 82 per cent, which Mistral says is the highest of any model. Then comes the sentence that matters:

Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.

Near zero is not incompetence. It is a refusal boundary, deliberately placed. Anthropic and OpenAI can presumably do this work and have decided their models should not. Mistral's argument for crossing that line is a defender's argument, and it is a reasonable one on its face:

Defending software often starts with proving that a flaw is real, exactly the kind of work safety filters in closed models can block.

That is true. Security teams genuinely do get blocked by refusals on legitimate work, and a model that scores 93 per cent on Cybench — forty exercises drawn from security competitions — is useful to people whose job is finding holes before someone else does. Co-founder Guillaume Lample framed it as sovereignty: cyber defence capability that lets enterprises and governments defend themselves.

The problem is the next step. A refusal boundary is a control that lives in the serving stack. It can be tuned, audited, logged and revoked. Once weights are downloadable, none of that survives. Asked about this, VP of science Pierre Stock told TechCrunch:

We'll work with trusted partners and governments to make sure that the open source weights can be used to defend, but not to perform malicious attacks.

There is no mechanism that does this. A weights release is not a partnership; it is a file. Mistral has not yet named the licence Large 4 will ship under, and a licence is in any case a legal instrument, not a technical one. The honest version of the position is that Mistral has decided the defensive value is worth the diffusion, and that it would rather Europe hold this capability than cede open-weight cyber tooling to Chinese labs. That is a defensible bet. It is not the same thing as a safeguard, and the three weeks between preview and weights are the window in which anyone who disagrees has to say so.

The rest of the numbers are Mistral's own

Everything below comes from Mistral's charts, measured before independent reproduction. Treat it accordingly.

BenchmarkLarge 4Named comparison
DeepSWE v1.1 (agentic coding)61.7%GLM-5.3 61%, DeepSeek-V4-Pro 57%
FinWorkBench (finance)67%DeepSeek-V4-Pro 67%
Harvey Legal Agent15%Kimi K3 13%, GPT-6 Astra 5%
DIOR-RSVG (visual grounding)73%GPT-6 Astra 68%
Terminal-Bench 4.028.3%—
Cybench93%—

VentureBeat noted the obvious caveat: benchmark configurations differ between vendors, and public leaderboards put GLM-5.3 nearer 69 per cent on DeepSWE than the 61 per cent Mistral's chart shows. Artificial Analysis, running its own harness, scored Large 4 at 38 on its Intelligence Index — enough, it said, to make France once again the home of the most capable model from outside the United States and China. That is the number worth holding onto, because it was not produced by Mistral.

The efficiency claim is similarly self-reported and similarly interesting. Mistral says Large 4 was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own European datacentres over roughly two months, drawing about ten megawatts. Stock put it comparatively: around 4,000 GPUs, "two to three times less than our Chinese competitors, and significantly less than the closed source competitors." If that holds up after the weights land, it is a more consequential result than any benchmark on the list — a frontier-class model trained on a cluster an order of magnitude smaller than the ones Meta and OpenAI talk about, by a company that had three researchers and now has about three hundred, valued at €21 billion after a €3 billion round last month.

What to watch

The licence, when it is named. Whether Artificial Analysis's Cyber Index score survives third-party reruns against the released weights. And whether any other lab follows — because if an open-weight model with top-five cyber capability is downloadable and nothing bad visibly follows, the refusal boundaries at Anthropic and OpenAI start to look like a competitive disadvantage rather than a safety position. That is the argument Mistral has just forced everyone else to have.

Sources: Mistral AI: Introducing Mistral Large 4, TechCrunch: Mistral's new 1T model aims to leapfrog closed and open rivals, VentureBeat: Mistral debuts Large 4 'Le Chonk', The Next Web: Europe's Mistral launches Large 4 to challenge China's lead in open AI models, MarkTechPost: Mistral AI releases Mistral Large 4 (Le Chonk), Artificial Analysis on X: Mistral Large 4 scores 38 on the Intelligence Index

Keep reading