AI/News

Mistral's Large 4 is a trillion-parameter model at $1.36 per million tokens in, with the weights public by the end of October - and it does the vulnerability work closed models refuse

Mistral's own page: 52 billion active parameters (the page read 49 billion on 6 and 7 October and was changed by Mistral - corrected here 8 October), trained on 3,800 Grace Blackwell GPUs in its own European datacentres, a preview API today at $1.36 in and $4.18 out per million tokens. On a test that asks a model to reproduce a real software flaw and patch it, ML4 scores 82% - while, Mistral says, Claude Opus 5.5 and GPT-6 Astra "score near zero on the same test because they refuse to perform the task".

By Daily Aletheia · Checked against the primary source · 7 October 2026 · 3 min read


Image: Mistral - the band from the company's own Mistral Large 4 announcement art (its pixel mascot, 'le Chonk'), mistral.ai, 6 October 2026

Mistral launched a public preview of Mistral Large 4 on 6 October - "Unofficially ML4, very officially: le Chonk" - and said the weights "drop end of this month". The page calls it "a 1 trillion-parameter natively multimodal model with 52 billion active parameters" (Correction, 8 October: when this story ran on 7 October the page read "49 billion active parameters" and we carried that figure; Mistral has since changed its page to 52 billion, and the copy now carries the page as it stands.), the company's "largest and most capable model to date", trained "from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters in Europe". The preview API is on Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens.

The claim. The model "already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe". On "critical enterprise workloads, including cybersecurity, finance and law", Mistral calls it "state-of-the-art among open models"; on visual grounding it says the model surpasses "even frontier closed models", giving Dense 200 at 42% against GPT-6 Astra's 41%.

The cyber numbers. On the Artificial Analysis Cyber Index the page says ML4 "ranks among the top five models globally and leads open-weight models developed outside China by a wide margin". On one of that index's tests, "which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model"; it "solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions". Then the line that matters for anyone choosing a model for security work: "Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task." Mistral's answer is open weights and self-deployment, "giving organizations both the capability and the autonomy to run advanced security work under their own policies". Until the weights ship, the page says, the model is being red-teamed "with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities".

Coding and agents, as the page reports them. 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4; a Coding Agent Index of 49.8%, which Mistral places "ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max". In a blind human evaluation run by Surge AI, "ML4 Preview ranked second of five models (3.74) ... and behind only Claude Opus 5 (4.22)". On AutomationBench, 657 business workflows across apps like Gmail, Google Sheets, Slack and Salesforce, 59.9%. On Lakera's B3 prompt-injection benchmark it "resists 93.3% of attacks".

What is still to come. Architecture details, further benchmarks and the post-training method arrive with the weights. The page ties the model to the "€3 billion Series D" Mistral closed earlier this year and says the reinforcement-learning run behind the preview "is still in flight", producing "roughly 33 billion tokens per day" at "3k GPUs".

The same question, answered the other way, the same day. Anthropic's page of 6 October opens three vetted tiers of cyber access to Claude and reports that, with no programme access, "every task was blocked on the first prompt" on its own cyber scenarios - Mistral's "near zero" put in Anthropic's words, on a different test. The open-weights route and the vetted-tier route are now both on the table for the same work; which one a security team can use will be the question of the month.

The Check

What was said
"Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task" - Mistral, about its competitors
What the Source Says
Mistral's figure is its own reading of a third-party index (Artificial Analysis) on one test; no score for the closed models is printed on the page. Anthropic's own page the same day reports, on a different evaluation (CyScenarioBench, 10 challenges, five attempts each), that without programme access "every task was blocked on the first prompt" for Claude Opus 5.5.
The Difference
The refusal is confirmed by Anthropic in its own words, but on a different test; Mistral's "near zero" for the named closed models is a vendor's claim about rivals and is carried here as that, not as a measured score.

Sources & further reading

  1. 01Mistral - Introducing Mistral Large 4, 6 Oct 2026 ↗Primary
  2. 02Anthropic - Expanding the Cyber Verification Program, 6 Oct 2026 (the other answer to the same question) ↗Primary

Corrections

  • 8 October 2026 - Active-parameter count: Mistral's own page read "49 billion active parameters" on 6 and 7 October (our saved copy of the page) and now reads "52 billion active parameters". Title unchanged; dek, body and the Truth Check row updated on 8 October to the page as it stands, with the earlier figure named in the story. Not a misreading - the source changed its page after we quoted it.