Read online version of the 7 October 2026 issue
Newsletter/7 October 2026
A trillion-parameter model for $1.36 a million tokens, open-weight by month-end, that does the cyber work Claude and GPT refuse. OpenAI's 722 maths papers: 4,000 problems in, not all checked.
Mistral's open weights against Anthropic's three vetted tiers - two answers, one day, to who gets the model that proves a flaw is real; "OpenAI's AI produced 722 new mathematical results" against the repository's own README; and the record check, one prompt that tells you whether your agent can edit its own history.
By Daily Aletheia · Checked against the primary source · 7 October 2026 · 5 min read

On Monday Mistral put a trillion-parameter model on a public API at $1.36 a million tokens, promised the weights by the end of the month, and said plainly what it does that Claude Opus 5.5 and GPT-6 Astra refuse to: reproduce a real software flaw and patch it. The same day Anthropic opened three vetted tiers of exactly that capability, with a form. And OpenAI published 722 mathematics papers written by a model it will not name.
The Signal - Mistral puts a trillion-parameter model on the table, weights free at month-end; Anthropic answers with three vetted tiers and a form
What happened: On 6 October Mistral launched a public preview of Mistral Large 4, which its page calls "a 1 trillion-parameter natively multimodal model with 52 billion active parameters", at $1.36 per million tokens in and $4.18 out, with the "weights drop end of this month". The same day Anthropic expanded its Cyber Verification Program into three tiers - Defense, Red Team and Specialized Access - that lift the cyber blocks on Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1.
The details:
- The capability in question: on a third-party test that asks a model to reproduce a real vulnerability and patch it, Mistral says ML4 scores 82%, "the highest of any model", while Claude Opus 5.5 and GPT-6 Astra "score near zero on the same test because they refuse to perform the task".
- Anthropic's page says the same in its own numbers: with no programme access Claude Opus 5.5 was blocked on "every task" of its ten cyber scenarios; in the Defense tier 46 of 50 trials were blocked; in the Red Team tier none, and it completed 34 of 50.
- The price of each route: Mistral's is the weights, "on private cloud or on-premise", "under their own policies". Anthropic's is a form - "a few days" for Defense, "a few weeks" for Red Team - with data retention "required ... so that we can monitor for cyber misuse", the top tier reviewed "in collaboration with the US government".
- Found so far: nothing from Mistral's preview; Anthropic's programme reports "at least 129,000 verified software vulnerabilities between April and July 2026".
What it means for you: The capability closed models refuse is about to be downloadable. Until the end of October the choice is a form; after it, a policy you write yourself. Either way the model that proves a flaw is real is the model that can exploit it; what changed this week is that one vendor will hand it over and one will watch you use it.
The counterweight: Every figure above is a vendor's own. Mistral's "near zero" for its rivals is its reading of someone else's index; Anthropic's 129,000 comes from "33 partner reports" and it calls the number an undercount.
Correction, 9 October: this paragraph gave Mistral Large 4 as having 49 billion active parameters - the figure on Mistral's page on 6 and 7 October, read verbatim. Mistral changed its page to "52 billion active parameters" on 8 October; the figure above now matches the page as it stands.
The Truth Check - "OpenAI's AI produced 722 new mathematical results"
Free to keep reading. Every issue, checked, three mornings a week.
Free. Unsubscribe any time.
The Move - The record check
The Move - The record check is for members. $9 a month, founding price.
The Check
- What was said
- "OpenAI's AI produced 722 new mathematical results" - OpenAI's own sentence ("a broad range of new mathematical results produced by an internal frontier model"), carried as 722 AI-generated proofs
- What the Source Says
- The repository README: "722 manuscripts organized into 372 families"; "approximately 4,000 problems" posed; "three hours of ChatGPT Pro thinking compute" per result on average; "Many, but not all, of the manuscripts have been formalized"; "This collection includes results at different stages of verification ... Some of the unformalized results could have issues."
- The Difference
- The count holds and the method is stated; "verified" or "proved" does not hold for the set - a paper without a Lean formalisation is a claim awaiting a referee, and OpenAI says so on the page. The model is unnamed and unreleased.
Sources & further reading
- 01Mistral - Introducing Mistral Large 4, 6 Oct 2026 ↗Primary
- 02Anthropic - Expanding the Cyber Verification Program, 6 Oct 2026 ↗Primary
- 03OpenAI - Sharing AI progress in mathematics, 6 Oct 2026 ↗Primary
- 04OpenAI - the openai/math repository and its README, 6 Oct 2026 ↗Primary
- 05Hacker News - front page, read 7 Oct 2026 (the two stories' points) ↗Attributed
- 06METR - AI systems could cover up misbehavior, 6 Oct 2026 ↗Primary
- 07BBC News - Pentagon stops using Anthropic AI tools after blacklisting company, BBC told, 5 Oct 2026 (read as carried verbatim, bylined BBC News, by The Star) ↗Attributed
- 08OpenAI - Building advertising for the way people use AI, 5 Oct 2026 ↗Primary
- 09OpenAI - Our approach to EU text provenance rules, 5 Oct 2026 ↗Primary
Corrections
- 9 October 2026 - Mistral Large 4 active parameters: 49 billion -> 52 billion. Mistral changed its own page on 8 October 2026 (it read 49 billion on 6 and 7 October, when the issue quoted it). The Signal now carries 52 billion with a dated note.
Was this issue useful?