Read online version of the 7 October 2026 issue

Newsletter/7 October 2026

A trillion-parameter model for $1.36 a million tokens, open-weight by month-end, that does the cyber work Claude and GPT refuse. OpenAI's 722 maths papers: 4,000 problems in, not all checked.

Mistral's open weights against Anthropic's three vetted tiers - two answers, one day, to who gets the model that proves a flaw is real; "OpenAI's AI produced 722 new mathematical results" against the repository's own README; and the record check, one prompt that tells you whether your agent can edit its own history.

By Daily Aletheia · Checked against the primary source · 7 October 2026 · 5 min read


Image: Mistral - the band from the company's own Mistral Large 4 announcement art (its pixel mascot, 'le Chonk'), mistral.ai, 6 October 2026

On Monday Mistral put a trillion-parameter model on a public API at $1.36 a million tokens, promised the weights by the end of the month, and said plainly what it does that Claude Opus 5.5 and GPT-6 Astra refuse to: reproduce a real software flaw and patch it. The same day Anthropic opened three vetted tiers of exactly that capability, with a form. And OpenAI published 722 mathematics papers written by a model it will not name.

The Signal - Mistral puts a trillion-parameter model on the table, weights free at month-end; Anthropic answers with three vetted tiers and a form

What happened: On 6 October Mistral launched a public preview of Mistral Large 4, which its page calls "a 1 trillion-parameter natively multimodal model with 52 billion active parameters", at $1.36 per million tokens in and $4.18 out, with the "weights drop end of this month". The same day Anthropic expanded its Cyber Verification Program into three tiers - Defense, Red Team and Specialized Access - that lift the cyber blocks on Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1.

The details:

  • The capability in question: on a third-party test that asks a model to reproduce a real vulnerability and patch it, Mistral says ML4 scores 82%, "the highest of any model", while Claude Opus 5.5 and GPT-6 Astra "score near zero on the same test because they refuse to perform the task".
  • Anthropic's page says the same in its own numbers: with no programme access Claude Opus 5.5 was blocked on "every task" of its ten cyber scenarios; in the Defense tier 46 of 50 trials were blocked; in the Red Team tier none, and it completed 34 of 50.
  • The price of each route: Mistral's is the weights, "on private cloud or on-premise", "under their own policies". Anthropic's is a form - "a few days" for Defense, "a few weeks" for Red Team - with data retention "required ... so that we can monitor for cyber misuse", the top tier reviewed "in collaboration with the US government".
  • Found so far: nothing from Mistral's preview; Anthropic's programme reports "at least 129,000 verified software vulnerabilities between April and July 2026".

What it means for you: The capability closed models refuse is about to be downloadable. Until the end of October the choice is a form; after it, a policy you write yourself. Either way the model that proves a flaw is real is the model that can exploit it; what changed this week is that one vendor will hand it over and one will watch you use it.

The counterweight: Every figure above is a vendor's own. Mistral's "near zero" for its rivals is its reading of someone else's index; Anthropic's 129,000 comes from "33 partner reports" and it calls the number an undercount.

Correction, 9 October: this paragraph gave Mistral Large 4 as having 49 billion active parameters - the figure on Mistral's page on 6 and 7 October, read verbatim. Mistral changed its page to "52 billion active parameters" on 8 October; the figure above now matches the page as it stands.

The Truth Check - "OpenAI's AI produced 722 new mathematical results"

Free to keep reading. Every issue, checked, three mornings a week.

Free. Unsubscribe any time.

The Move - The record check

The Move - The record check is for members. $9 a month, founding price.

The Check

What was said
"OpenAI's AI produced 722 new mathematical results" - OpenAI's own sentence ("a broad range of new mathematical results produced by an internal frontier model"), carried as 722 AI-generated proofs
What the Source Says
The repository README: "722 manuscripts organized into 372 families"; "approximately 4,000 problems" posed; "three hours of ChatGPT Pro thinking compute" per result on average; "Many, but not all, of the manuscripts have been formalized"; "This collection includes results at different stages of verification ... Some of the unformalized results could have issues."
The Difference
The count holds and the method is stated; "verified" or "proved" does not hold for the set - a paper without a Lean formalisation is a claim awaiting a referee, and OpenAI says so on the page. The model is unnamed and unreleased.

Sources & further reading

  1. 01Mistral - Introducing Mistral Large 4, 6 Oct 2026 ↗Primary
  2. 02Anthropic - Expanding the Cyber Verification Program, 6 Oct 2026 ↗Primary
  3. 03OpenAI - Sharing AI progress in mathematics, 6 Oct 2026 ↗Primary
  4. 04OpenAI - the openai/math repository and its README, 6 Oct 2026 ↗Primary
  5. 05Hacker News - front page, read 7 Oct 2026 (the two stories' points) ↗Attributed
  6. 06METR - AI systems could cover up misbehavior, 6 Oct 2026 ↗Primary
  7. 07BBC News - Pentagon stops using Anthropic AI tools after blacklisting company, BBC told, 5 Oct 2026 (read as carried verbatim, bylined BBC News, by The Star) ↗Attributed
  8. 08OpenAI - Building advertising for the way people use AI, 5 Oct 2026 ↗Primary
  9. 09OpenAI - Our approach to EU text provenance rules, 5 Oct 2026 ↗Primary

Corrections

  • 9 October 2026 - Mistral Large 4 active parameters: 49 billion -> 52 billion. Mistral changed its own page on 8 October 2026 (it read 49 billion on 6 and 7 October, when the issue quoted it). The Signal now carries 52 billion with a dated note.

Was this issue useful?