358 Years. 11 Days. What Claude actually did to Fermat

Claude Didn't Prove Fermat's Last Theorem. What It Actually Did Is More Interesting.

Anthropic's Claude wrote 13 million lines of Lean code in 11 days. The mathematician paid £1 million to do this by hand found out at a music festival and assumed it was a crank. Here's what actually happened — and why most headlines got it wrong.

TL;DR

- On September 4, 2026, Anthropic announced that Claude produced the first complete, computer-checked formalization of Fermat's Last Theorem in 11 days, writing 13 million lines of Lean code - This is not a new proof — Andrew Wiles already proved it in 1995. What Claude did was translate that human proof into a form a computer can verify line by line - Kevin Buzzard, the mathematician paid £1 million over five years to do this same task by hand, downloaded Claude's code, compiled it himself, and wrote on his blog: "Anthropic has beaten me to it" - The first attempt failed completely — agents lost track of what they were doing. The breakthrough came not from a better model but from better coordination software - The 13 million lines of Lean code can't enter Mathlib yet, no new mathematics was discovered, and the result produces nothing a human can read. Buzzard says this "tells us essentially nothing" about mathematics — and also that it changes everything

There's a version of this story that writes itself: AI conquers 358-year-old math problem that stumped every human who ever lived. Triumphant. Clean. Wrong.

The accurate version is stranger and more interesting. Here's what actually happened.

The thing most headlines missed

Fermat's Last Theorem was proved in 1995. Andrew Wiles, a Princeton mathematician who spent seven years working in near-total secrecy, published a 129-page proof that took months of expert review to verify. Most number theorists today put the probability it contains an error at effectively zero.

What Claude produced is something different: a formalization of that proof. The difference matters.

A human proof is written for other humans. It leaves enormous amounts implicit — standard constructions, obvious steps, things the reader can fill in because they've seen the pattern before. Mathematicians are comfortable with this. A proof assistant like Lean is not. Lean demands that every single logical step be spelled out explicitly. No gaps, no "the reader can verify," no trust. If a single step is missing or wrong, the whole thing fails to compile.

Turning Wiles's 129 pages into something Lean accepts is an enormous undertaking. Dutch computer scientist Jan Bergstra proposed doing it in 2005. Twenty-one years later, it still wasn't done. Kevin Buzzard of Imperial College London kicked off a community effort in 2024, backed by a £1 million grant from the UK's EPSRC, with a five-year timeline.

Anthropic took eleven days.

A mathematician gets an email at a music festival

Buzzard heard about the result while at a music festival with poor reception. An email arrived from a stranger with the subject line "End-to-end Lean formalization of Fermat's Last Theorem." He assumed it was a crank.

It was not a crank.

When he got back, Buzzard did what any rigorous mathematician would do: he downloaded the repository, compiled it himself on a 96-core machine, and ran the kernel checker. The proof passed. He checked specifically for cheating — Lean has had soundness bugs found recently, and a clever agent could in theory exploit one to "prove" anything. He asked an agent to flag every line that wasn't a definition or a proof. About 100 lines came back; they defined a convenience tactic. OpenAI's models also reviewed the Lean codebase and found no soundness issues in the version Claude used.

His blog post was titled: "FLT: Anthropic has beaten me to it."

His verdict, characteristically honest, came in two parts. On the mathematics: it changes nothing. The formalization "just faithfully follows the early literature on the proof and adds nothing." On what it demonstrates: everything. An AI swarm just formalized thousands of pages of literature in eleven days. If that's possible, formalization of modern research could start happening as a matter of routine.

He added one sentence no press release would carry. He was given £1 million to run his project over five years. He wondered whether Anthropic spent more.

The arithmetic on that wondering

Anthropic says the run consumed about six billion output tokens. The model was an internal research model comparable to Claude Fable 5.1, which costs $50 per million output tokens at list price. Six billion tokens at that rate is $300,000.

That's an illustration, not an invoice. The model was internal, inference costs less than list price, and nobody raised an actual bill. But it's the only public arithmetic available, and it sits against £1 million spread over five years.

The scale of the thing

13 million lines of Lean code. 29,500 intermediate theorems proved (30,300 produced; 29,500 used). The resulting proof is more than five times the size of Mathlib, the community's entire mathematics library that hundreds of researchers have built over years.

Compilation on that 96-core machine took nearly twenty times longer than Mathlib itself.

Tianyi Peng, the Anthropic researcher at Columbia who ran the project, gave the agents essentially two sentences of guidance: "Jacobian as a scheme sounds high priority." "Push Mazur to be done soon." The rest was autonomous.

The proof follows the 1995 exposition of the Wiles argument by Darmon, Diamond, and Taylor — not the modern route Buzzard's team was working on. It uses only Lean's three standard axioms (propext, Classical.choice, Quot.sound). A comparator confirmed the theorem statement matches Mathlib's own definition of Fermat's Last Theorem. No assumptions other than the axioms of mathematics.

It also, quietly, closes Freek Wiedijk's "100 theorems" list — a 20-year benchmark project tracking the formalization of the 100 most important mathematical results. FLT was the last one.

The first attempt failed

This detail is the most instructive thing in Anthropic's research post, and most coverage skipped it.

The initial runs didn't work. Agents made early progress, then lost track of the project's state and stopped coordinating effectively. The effort collapsed. Those failed runs weren't entirely wasted — they contributed about 7% of the non-boilerplate lines in the final proof — but they didn't produce a proof.

What fixed it wasn't a better model. The weights didn't change.

What changed was scaffolding: a platform called Prove2Me, built by Peng with collaborators at Columbia. It maintains a directed graph of theorem statements so agents can see what to attempt next. It splits statements and proofs into separate files to speed compilation. It keeps a plain-language description of each statement so work can be found and reused across agents.

Add a multi-agent harness built on Claude Code, and the same models that previously failed finished the job in under a fortnight. Same capabilities, different coordination. The scaffolding decided the outcome.

What this doesn't mean

Buzzard is careful about what the result shows, and it's worth following his reasoning rather than the headlines.

No new mathematics was created. The formalization follows Wiles; Claude didn't find a new route, didn't discover a shortcut, didn't extend the result. In the context of ongoing AI mathematics claims — OpenAI said in August that Astra had solved ten open problems; researchers found flaws in that methodology — a formalization is a different and more modest object. It's verification, not discovery. Discovery is harder.

None of the 13 million lines can currently enter Mathlib. Mathlib's maintainers don't yet accept AI-generated code submissions, partly because the review backlog is already over 3,000 open pull requests and reviewers are reluctant to work through AI output. The constraint hasn't disappeared; it's shifted from "writing proofs" to "reading them."

The proof also doesn't produce anything a human mathematician can learn from. There's no human-readable exposition, no document explaining the mathematical ideas in a way that builds intuition. Buzzard's own grant committed him to building that document. He doubts Anthropic will do it.

What this does mean

Kevin Buzzard — the person with the most expertise and the most reason to find a flaw — put it precisely: if the automatic formalization of FLT is possible now, we've taken a big step toward automatic formalization of the modern mathematical literature.

The implications run in several directions.

For mathematics: as AI produces more purported proofs, formal verification becomes the natural quality check. Instead of a result sitting in a journal and taking months or years for the community to fully verify, a formalization could accompany the original write-up. Errors that currently persist in the literature for years because nobody checks every step would surface immediately.

For AI: this is a cleaner capability signal than a leaderboard number. The proof either compiles or it doesn't. It's either sorry-free or it isn't. Buzzard either signs off on it or he doesn't. In a week where benchmark scores were collapsing 36 points under independent testing, a binary artifact that the rival researcher compiles himself is a different class of evidence. Anthropic chose to publish an artifact; other labs chose to publish scores. That asymmetry will matter more as capability claims multiply.

For everyone building AI agents: the lesson from the failed first attempts is worth sitting with. Better coordination architecture — not a better underlying model — was the variable that determined success. Prove2Me is open. The multi-agent harness is built on Claude Code. The thing that made the difference is replicable by anyone paying attention.

Three researchers on personal Claude Max subscriptions also formalized the Hardy-Littlewood Circle Method — a technique central to analytic number theory — in three days, as a side experiment. It worked. The approach generalizes.

Buzzard still has a job. His grant committed him to things Claude didn't do: submitting foundational number theory to Mathlib, and building the human-readable document that lets mathematicians actually understand the modern proof. He doubts the code churned out at $300,000 over eleven days will do the second part.

He's probably right. And that's the most useful frame for this whole story: the AI did the mechanical part, faster and cheaper than the humans. The humans still own the meaning.