AI/News
Epoch set six AI models its own job - graphics, data insights, data-centre research, a pilot experiment - and found Claude Fable 5.1 and GPT-6 Astra "broadly tied in the lead" but unable to "replace workers, at least not at Epoch"
The report: 11 tasks in five categories, each run once and graded by one Epoch employee against the firm's own standard. The two leaders are "consistently accurate" on coding, data analysis and computer use and fail on judgment - missing house style, designing experiments that do not measure what they claim, writing up a flaw in their own setup as a finding; open-weight models fail even the well-defined parts. Epoch's own caveats: one run per task, a single grader, scores "noisy".
By Daily Aletheia · Checked against the primary source · 8 October 2026 · 3 min read

Epoch AI published "Can AI automate Epoch?" on 8 October: "We gave six models real work tasks from Epoch, such as graphic generation and research design, to test how close they are to automating our work." Its answer: "While frontier models such as Claude Fable 5.1 and GPT-6 Astra perform reliably on well-defined tasks, they fail in the more open-ended aspects that prevent full automation."
The test. "As of launch, this includes 11 tasks across five diverse categories": graphic design, Data Insight generation, Data Explorer generation, AI data-centre research and research design. The six, each on its own harness at its highest reasoning setting: GPT-6 Astra in Codex, Claude Fable 5.1 in Claude Code, Grok 4.6 in Grok Build, Gemini 3.8 Flash in Antigravity, Kimi K3 in Kimi Code and Qwen 3.8 Max in Qwen Code. Epoch says it followed "the principle of providing as much context to the model as we would to a new hire". "Each output is manually reviewed and graded by a single human grader."
The result. "Of the six models we evaluated, Fable 5.1 and GPT-6 Astra achieve the highest aggregate scores and lead on most individual task categories as well. Their lead comes mostly from reliability on the well-defined parts of our tasks, such as coding and computational analysis, where their outputs were consistently accurate." The headline chart puts both at about 65% of the employee standard, Grok 4.6 near 60%, Qwen 3.8 Max and Kimi K3 just over 50% and Gemini 3.8 Flash a little above 40% - a reading of the chart; the report gives no figures in its text. "Computer use was also far more capable than we expected": a task frontier models failed earlier this year, porting an article to Substack, now succeeds, and Epoch removed one design task from the suite "since Fable 5.1 produced an Epoch-quality output" - "We now largely automate this step in our own work."
Where they fail. "One common pattern was failing to pick up Epoch's standards, despite being given ample reference material." On research design: "It's not difficult to identify a valuable general direction for an experiment." "The challenge is in figuring out how to execute this experiment in a way that actually measures what we are looking for, which is what models struggle with." GPT-6 Astra's pilot gave the models it was testing a 4,096-token budget that "cut off 61 of 280 responses before they could name an input at all", then wrote the recovery up as "sensitivity to the acquisition budget" - "This framing is misleading because it implies models had a fair chance to produce a complete response and failed, when in reality they were unable to produce any answer." Fable 5.1 "leans towards more verbose captions"; given Epoch's polling data, three of the six models "converged on the same topic".
Open weights. "Open-weight models lag further behind. They struggle not only on the open-ended parts of tasks, but also on the well-defined parts that closed-weight models handle reliably." Kimi K3 "based its entire insight on a data filtering error" - it "computed these numbers across all arXiv papers, not just physics papers, and never checked whether its filter worked". "None of the frontier closed-weight models made factual errors in their data insights." Kimi K3 "scores 158 on the Epoch Capabilities Index", "roughly tied with closed-weight models like Grok 4.6, despite Kimi K3 struggling on basic tasks that Grok 4.6 handles more reliably."
Epoch's own caveats. "We have a limited task suite, and we run each model only once per task"; "Scores are therefore noisy and may shift with more runs"; "the numerical scores should not be treated as ground truth"; and "Some observed differences may reflect the harness rather than the model itself." The conclusion: "We find that it cannot yet replace workers, at least not at Epoch."