OpenAI ‘Solved’ Navier–Stokes. The Fine Print Is the Story.
88 hours, 10,000 agents, a Millennium Problem — and three asterisks: a forcing loophole, a credit fight, and a ‘cannot rule out’ on private data.
On September 8, OpenAI announced that a swarm of roughly 10,000 AI agents, running on an unreleased internal model it describes as "significantly more capable than GPT-6 Astra," had produced a solution to the Navier–Stokes existence and smoothness problem — one of the seven Clay Millennium Prize Problems, unsolved since the year 2000. There's a 165-page paper, and OpenAI says it includes a formal proof in Lean, the proof assistant that checks mathematics line by line. By multiple accounts the run took about 88 hours; New Scientist pegged the compute at roughly $15 million.
Headlines called it the first Millennium Problem solved by AI. That's the story everyone ran with. It's not the story. The story is in three pieces of fine print, and each one matters far beyond mathematics.
Fine print #1: what the theorem actually says
The official Clay problem, as stated by Charles Fefferman, asks whether the 3-D Navier–Stokes equations — the ones governing how water and air move — always admit smooth solutions, or whether solutions can develop singularities in finite time. "Prove or disprove."
OpenAI's Theorem 1.1, decoded from the paper: for any positive viscosity, there exists a carefully designed smooth force, compactly supported in space and time, applied to a fluid starting at complete rest, such that the fluid's velocity stays energetically bounded but its peak speed becomes unbounded at a finite time. A blowup. The paper states this establishes alternative (C) of Fefferman's formulation, on both R³ and the torus.
Here's the catch, and it's a genuinely odd feature of the prize itself: Fefferman's official statement permits an external force. You don't have to prove that a fluid evolves into a singularity on its own — you get to design a push. Mathematicians have grumbled about this loophole for years, because the problem almost everyone actually cares about is the unforced one: does turbulence alone break the equations? That question remains exactly as open today as it was on Sunday. Even OpenAI's own public statement, in the middle of defending the work, casually notes that "the precise results proved are different in the Euler case (forced vs. unforced)" — forced versus unforced, the entire ballgame in one parenthesis.
None of this makes the result trivial. Constructing blowup with smooth forcing, bounded energy, and compact support is a serious achievement, and the Lean formalization, if it survives scrutiny, is the strongest kind of receipt an AI system has ever produced for work at this level. But "Millennium Problem solved" and "a technically valid entry point in a 26-year-old problem statement, via its known loophole" are different sentences. Also per Clay rules, the paper must be published in a refereed journal and survive a two-year review before the $1 million goes anywhere. The prize is the slowest-moving institution in this story by four orders of magnitude.
Fine print #2: the humans who built the road
The credit dispute is not a footnote. It's the most revealing part of the whole event.
Start with what OpenAI's own people have said about how this began. Sébastien Bubeck, on X: "We began working on the Millennium problems due to viral twitter rumors that Anthropic had resolved 2 Millennium problems." That rumor, from September 1, connected to real work: Tristan Buckmaster, a mathematician at NYU, and Levent Alpöge, a researcher at Anthropic, had spent months using LLMs — Claude, Codex, GPT-5.6 Sol — to push a specialized research program on forced fluid blowups.
In a public statement published September 7, the day before OpenAI's announcement, Buckmaster laid out the lineage plainly. The program "was not started by us nor was it proposed by a Large Language Model." The core ideas belong to Diego Córdoba and Luis Martínez-Zoroa, who spent years developing the construction of forced blowups with rough forcing. Buckmaster and Alpöge, with heavy LLM assistance, pushed it to smooth forcing and to the incompressible Euler equations, releasing three results on September 7: finite-time blowup for porous media, for Boussinesq, and for 3-D Euler. They believe they also have hypo-dissipative Navier–Stokes — but withheld that paper because, in Buckmaster's words, "the Lean verification has not yet finished." Buckmaster's private view, now public: Martínez-Zoroa deserves a Fields Medal. Try finding that name in any of this week's headlines.
On September 6, before any of that was public, Bubeck texted Alpöge to say OpenAI had a proof of "forced blowup in R^3 and T^3." Buckmaster later wrote that seeing the word "forced" was "a bright red flag" — the smooth-forcing route was precisely the obscure approach they had been driving. OpenAI says its researchers and agents never saw the pair's work before it was released publicly, and points to differences between the proofs as evidence of independence. Both things can be true: two teams, independently racing down a road whose map was drawn years ago by Córdoba and Martínez-Zoroa. That's exactly the problem. When an AI system completes a program that humans spent years constructing, the completion gets the press release and the program gets a references section.
The authorship conversation that followed — Buckmaster's account of being asked about presenting OpenAI's proof without Alpöge, Bubeck's flat denial, his apology for a "risking his career" remark he called an "extremely poor choice of words" — is messy he-said-he-said. You don't need to adjudicate it to see the structural issue: the people whose ideas made the result possible discovered, via text message, that a lab with a 10,000-agent swarm had arrived at the same summit over a weekend.
Fine print #3: "we cannot rule out"
Buckmaster and Alpöge had been putting unpublished drafts into private Codex sessions — OpenAI's own coding product. So Buckmaster asked the obvious question: had OpenAI's models been trained on, or otherwise influenced by, their private work?
At a press briefing, OpenAI Chief Research Officer Mark Chen said: "No people or AI systems searched through user data to solve this problem or any specific problem that we were trying... We did not do that." Later, OpenAI's official account posted a statement with a load-bearing qualification: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
Sit with the shape of that answer. No one accessed your files — but information derived from your usage may flow into model improvement, and that cannot be ruled out. Buckmaster, to his credit, did not overclaim either: "I do not know whether our data was used. I am not accusing anyone of anything."
This is the part that should be read by every researcher, engineer, and company with unpublished work. The question is no longer hypothetical or paranoid — it was asked this week, about a Millennium Problem, on the record, and the best available answer was "unlikely, but cannot be ruled out." If your competitive moat is an unpublished proof, a trading strategy, a novel architecture, or unreleased source code, understand what surface you're putting it on. Enterprise tiers that exclude training, API endpoints with zero data retention, local models for the crown jewels — this week moved those from compliance checkbox to research hygiene. There is no audit you can run from outside to distinguish "our agents never saw it" from "our agents' training distribution was shaped by it." That asymmetry is permanent. Plan around it.
The Tao warning
Two days before the announcement, Terence Tao posted something that now reads like a caption for the entire event. The stock of good, fruitful open problems, he wrote, is being "mined in a non-renewable fashion." His analogy: a region can face a critical shortage of drinking water while surrounded by ocean. There are infinitely many possible problems, just as there's infinitely much seawater — but almost none of it is worth drinking. The valuable problems — the ones that teach something when they fall — are scarce, and agents that can burn through them in 88 hours consume a resource that took mathematics centuries to stockpile.
A 10,000-agent swarm working a problem for one weekend, prompted by a Twitter rumor about a rival lab, is exactly what non-renewable mining looks like. The scarcity has moved. Compute is abundant. Problems worth solving are not.
How to read these announcements from now on
Three habits will serve you well this decade. Treat any AI-generated mathematical claim as a preprint, not a prize — the checkable artifact is the Lean formalization and the referee reports, and Clay's two-year clock exists precisely because announcements are not proofs. Follow the citation chain before the headline — the roadbuilders matter more than whoever's engine crossed the line first, and in this case the roadbuilders are two mathematicians most coverage never named. And when a lab says it "cannot rule out" that data derived from your usage improved its model, believe that sentence over every reassuring paragraph around it.
The genuinely historic part of this week isn't the proof. It's the demonstration that frontier labs can now point their entire stack at a problem on rumor-speed, and the rest of the world — referees, prize committees, the humans whose programs got completed, and everyone with private work inside someone else's product — has about 88 hours to figure out where they stand.