Truth Check/8 October 2026
“Epoch Automation Reports (8 Oct 2026): six models on 11 Epoch work tasks in five categories, one run each, one grader; Claude Fable 5.1 and GPT-6 Astra 'broadly tied in the lead' at about 65% of the employee standard (chart reading), open-weight models behind on the well-defined parts too; 'it cannot yet replace workers, at least not at Epoch'”
The claim holds against the page that produced it.
Epoch's own report page states the tie, the method (11 tasks, one run, a single grader) and the conclusion in its text; the 65% is read off its headline chart, which carries no figure in the text.
- Checked Against
- epoch.ai ↗Primary
- Read the Check
- Epoch set six AI models its own job - graphics, data insights, data-centre research, a pilot experiment - and found Claude Fable 5.1 and GPT-6 Astra "broadly tied in the lead" but unable to "replace workers, at least not at Epoch"
The report: 11 tasks in five categories, each run once and graded by one Epoch employee against the firm's own standard. The two leaders are "consistently accurate" on coding, data analysis and computer use and fail on judgment - missing house style, designing experiments that do not measure what they claim, writing up a flaw in their own setup as a finding; open-weight models fail even the well-defined parts. Epoch's own caveats: one run per task, a single grader, scores "noisy".
Every claim we have checked, in the Ledger.
Every claim checked against the page that produced it.
Three mornings a week, free.