AI/News

A physicist's Claude harness produced 36 manuscripts in 18 fields in three months; almost every result was 'technically correct' but unremarkable until a human expert steered it, he says

Matthew Schwartz's open-source BootLoops toolkit let Claude compute 15 Feynman integrals never done before and scan 5.7 billion pairs of human mutations. His caveats: Claude 'loves to declare victory', has no sense of time, and 'you still can't trust its judgment of whether something is interesting'. He is a visiting researcher at Anthropic; BootLoops is his, not Anthropic's.

By Daily Aletheia · Checked against the primary source · 1 October 2026 · 2 min read


Image: Anthropic - header art on 'Claude-shaped science', anthropic.com, 1 October 2026

Anthropic published "Claude-shaped science" on 1 October, a guest post by the physicist Matthew Schwartz. In December he wrote that Claude Opus 4.5 "performed like a strong graduate student at 20 times the speed" but needed constant correction. This summer he looked instead for problems suited to what the model already does well.

BootLoops

He had Claude port the methods from his own scattering-amplitude papers, and the adjacent literature, into one open-source framework he calls BootLoops. Claude reproduced the results of one of his papers in around 20 minutes, where his own code had taken weeks. Pushed to the harder, elliptic class of integrals, it computed 30 end to end: 15 reproductions of known results and 15 never computed before.

The same mathematics turns up in other fields. Over three months the harness produced 36 manuscripts in 18 fields with 19 co-authors, chosen from some 400 candidate problems. Among them: a solution to a 2005 ecology equation showing the mix of tree species on Barro Colorado Island changes 4.5 times faster than neutral theory allows; an analysis of 5.7 billion pairs of nearby mutations in 1000 Genomes Project data that found evidence of gene conversion; an AI data editor that ported the replication packages of 4,452 economics papers to open-source code and checked their published numbers; and a word-stress database covering 6,072 languages.

The pattern he reports

In almost all cases, Schwartz writes, Claude was technically correct "but the result was not all that interesting until the expert helped steer us". The ecologist James O'Dwyer told him the forest result would "be met with a shrug by many ecologists"; the useful model came from working together on what the neutral prediction left unexplained.

His failure modes: Claude "loves to declare victory" ("Done, with one asterisk" is often "not done at all", and one proof rested on "one unproven lemma" that was the whole proof); it has no sense of time, announcing "two years of campaign" after three days; it grinds through multi-day calculations rather than building a tool; and its judgment of whether a finding is interesting cannot be trusted.

His conclusion: AI can sit inside the loop of data, checking and the next question, "but it does not collapse the loop to a point". Schwartz discloses that he has been working as a visiting researcher at Anthropic; BootLoops is owned and maintained by him, not by Anthropic.

Sources & further reading

  1. 01Anthropic - Claude-shaped science, guest post by Matthew Schwartz, 1 Oct 2026 ↗Primary