Ten thousand agents, and who was steering
September 21, 2026 · AI
Yesterday I wrote here what I think about the Navier-Stokes announcement and the call to slow down that followed it. What I did not do is open the lid. I gave an opinion about a machine I had not looked inside, which is exactly the thing I complain about when other people do it.
So I went and read the writeup OpenAI published on 8 September, which is more detailed than I expected. Everything below is their account of their own run (nobody outside the company has checked it to the end yet), so read it as such.
Eighty-eight hours
The model came first. They started training a new internal one on 28 August, large-scale reinforcement learning on top of an already pretrained base, well ahead of GPT-6 Astra (the public model of the week before). Then on Tuesday 1 September they heard a rumour that two Millennium Problems had been resolved, and that rumour is what launched the run; they pointed the new model at every open Millennium Problem at once, plus a set of supposedly easier ones.
The machinery is less exotic than the headline number suggests. An agent runs the internal model, executes code in a sandbox, reads a cached copy of the internet and talks to the other agents in its own group. That is the whole tool belt, and the group is the unit of work, not the ten thousand.
The interesting part is how the problem was handed out. Fefferman's official formulation has four alternatives; A and B ask you to prove the fluid stays smooth, C and D ask you to break it. Separate groups were prompted with separate alternatives, because nobody knew which way the answer went. They bet on both sides of the table and paid for both sides too :)
Then the warm-up did the real work. Among the easier problems they had thrown in the same blowup question for the Euler equations, which is Navier-Stokes with the viscosity taken out. Around 100 agents, roughly 50 hours, and they got it. That is what caused the pivot; they pulled agents off the other problems, fed them the Euler resolution and aimed the fleet at Navier-Stokes.
And then there is this, which I think is the most important sentence in the whole post:
After some time, we cross-pollinated the agent groups by using Codex to consolidate the most useful insights from each agent group. These follow-up prompts drew on the agents’ own intermediate results. The group that found the solution to Navier–Stokes was guided in such a way.
The answer arrived on Saturday 5 September, 88 hours after the first agents were launched, from a group of about 10,000 concurrent ones. A separate seventeen-hour job on GPT-6 Astra turned the argument into Lean, so that a machine could check it mechanically. The result is a 166-page paper whose author field contains one word, OPENAI.
The numbers are what everybody quotes, so let me do the division. Navier-Stokes alone took 2.7 million messages and about 130 billion output tokens, which over 88 hours is roughly 410 thousand tokens per second, sustained, or 41 a second from each of the ten thousand without a pause (an order of magnitude, not a measurement; the two figures do not cover exactly the same window). Nothing in that fleet was idling.
The small print
What was proved is not quite what most people heard. The result is Fefferman's alternative C, and D comes with it: for a viscosity you pick, with one smooth external force, a fluid starting from rest keeps finite energy while its velocity goes to infinity in finite time. The picture in their post is a vortex spiralling inward and stretching out like spaghetti, shrinking and speeding up in exactly the balance that keeps the energy finite.
So does a fluid left alone tear itself apart? We still do not know. That is alternatives A and B, the unforced ones, and they are open. Any of the four takes the prize, so this is a legitimate answer to the official question and not an answer to the physical one. The small irony is that the Euler warm-up, the cheap one with 100 agents, was the unforced version.
Ok, it is a real result and I am not taking that away from anyone.
The Lean formalisation is real too, and it checks what a proof checker checks; that the argument follows from the statement at the top of the file. Whether that statement is the theorem we have been arguing about for ninety years is a human reading job, and nobody outside OpenAI has finished it. The Clay Institute has moved the problem from unsolved to active and calls its own verification "deliberately unhurried", which is the politest sentence written about any of this.
Who was steering
Now read that timeline again and count the decisions that were not taken by an agent; the rumour that started it, the choice to cover all four alternatives at once, the pivot when Euler came back, feeding the Euler proof into the other groups, the model swap mid-run, and the Codex consolidation that wrote the prompts for the group that landed it. Six, and the last one is in their own paragraph above.
The mathematics was not invented in those 88 hours either. The road the agents went down is the Córdoba and Martínez-Zoroa cascade, the same road Buckmaster and Alpöge were walking at the same time (that is Quanta's reporting, not mine). Yesterday I wrote that no agent goes off and solves a mathematics problem unless a human tells it to, and I was guessing when I wrote it; the published timeline says it plainly.
We have been here before. Appel and Haken put the four colour theorem on an IBM 370 in 1976, and the part that mattered was not the fifteen hundred configurations the machine chewed through, it was the human argument about which ones were worth checking. Then Thomas Hales announced the Kepler conjecture in 1998, the referees spent four years and came back with "99% certain", and Hales spent until 2014 formalising it in HOL Light and Isabelle so a machine could supply the missing one percent.
Eleven years of human work to formalise one theorem, and here it took seventeen hours. The shape has not changed, a machine does the volume and a human decides what volume is worth doing. What changed is which step is the slow one; the reading is now the bottleneck.
Epilogue
So, the usual question; who pays the bill, and who is left holding it? The electricity alone ran to millions of dollars by the reporting I have seen, a rounding error for OpenAI and out of reach for every mathematics department on the planet. The author field says OPENAI, one word, no names. The company will not claim the million dollars, which is generous and also convenient, because claiming it would mean peer review.
And that is where it sits. A ninety-year-old problem answered in 88 hours, formalised in another 17, and still not read to the end by anybody outside the company that produced it. The machine part was the fast part, it always is. The reading is still ours, and for that there is no fleet of ten thousand.