Truth Check/7 October 2026

“Claude Opus 5.5 and GPT-6 Astra score near zero on the same test because they refuse to perform the task”

The claim as it travelled: Mistral, 6 Oct 2026, about rivals' scores on a third-party index.

Unverified

No primary page supports the claim as stated.

Mistral prints no score for the closed models; the 'near zero' is its reading of a third-party index about its rivals, so it stands as Mistral's claim, not a measured figure.

Checked Against
mistral.ai ↗Primary
Read the Check
Mistral's Large 4 is a trillion-parameter model at $1.36 per million tokens in, with the weights public by the end of October - and it does the vulnerability work closed models refuse

Mistral's own page: 52 billion active parameters (the page read 49 billion on 6 and 7 October and was changed by Mistral - corrected here 8 October), trained on 3,800 Grace Blackwell GPUs in its own European datacentres, a preview API today at $1.36 in and $4.18 out per million tokens. On a test that asks a model to reproduce a real software flaw and patch it, ML4 scores 82% - while, Mistral says, Claude Opus 5.5 and GPT-6 Astra "score near zero on the same test because they refuse to perform the task".

Every claim we have checked, in the Ledger.

Every claim checked against the page that produced it.

Three mornings a week, free.