OpenAI's internal model cracked the Navier-Stokes Millennium Prize problem in 88 hours, according to the company's September 6 report "Research Acceleration: The View Inside OpenAI," a result that landed alongside a bitter public dispute over credit and conduct between the lab and mathematicians Tristan Buckmaster and Levent Alpoge.
The model, which is still training and had been training for less than two weeks, solved the problem 8 days after training started. It is not especially aimed at mathematics. Astra, the previous internal model, handled Lean formalization and verification in 17 hours.
OpenAI said it spent millions of dollars in inference to explore all unsolved Millennium Problems after hearing a false rumor on September 1, 2026, that two Millennium Problems had been resolved. Across all attempted problems, agents sent 4.9 million messages and used about 300 billion output tokens. For Navier-Stokes specifically, agents sent 2.7 million messages and used approximately 130 billion output tokens.
Depending on which costs are counted, the Navier-Stokes run would have cost a regular customer on the order of $22 million. Internal marginal costs were several million dollars. OpenAI does not intend to claim the Millennium Prize money. Sam Altman joked that AI is a bubble and that tokens are sold at a loss, worth only $1 million.
The result arrived with a warning attached. OpenAI and Anthropic have entered the RSI era, according to the company's own framing. The report says the bulk of the post is a detailed snapshot of how much agentic systems have contributed to RSI progress at OpenAI in recent months, and the answer is quite a lot. The surprise, the report says, is that use remained this low for so long. The report says the company is not not bragging and not not confessing. It is warning.
The Dispute Over Credit
Buckmaster and Alpoge worked for a year on blowup results. They achieved finite-time blowup with smooth forcing for incompressible porous media, Boussinesq, and 3D incompressible Euler. They believe they have a blowup for hypo-dissipative Navier-Stokes, but Lean verification is unfinished.
Buckmaster credits Diego Cordoba and Luis Martinez-Zoroa for the basic idea of the forced blowups program. He says Martinez-Zoroa deserves a Fields Medal. He apologized for early publication and compared the Euler writeup to AI slop.
Buckmaster reached out to OpenAI after hearing rumors. According to his account, OpenAI researcher Sebastien Bubeck said an internal OpenAI model had produced a proof of finite time blowup for forced Navier-Stokes over the past few days. Buckmaster claims that over two calls it became clear he was initially misled and that an entire team worked on the problem using "insane" compute.
Bubeck proposed two options, according to Buckmaster: Buckmaster posts the Euler result and OpenAI posts Navier-Stokes the next day, or Buckmaster alone writes a paper presenting the Navier-Stokes result while acknowledging that an OpenAI model resolved it. Bubeck twice asserted he wanted Levent removed from authorship, Buckmaster says, citing annoyance that Levent works at Anthropic. Bubeck said that if OpenAI posted after them, they would say Buckmaster and Alpoge deserved the Clay Prize and were "closest humans to the problem."
Buckmaster declined both offers and said he would go public if OpenAI released the result as proposed. Bubeck replied, according to Buckmaster, "Why would you ruin your career?" and "If you don't want me to be nice, then I don't have to be nice."
Buckmaster clarifies that he has not seen OpenAI's proof, does not know what their model did, does not know if their data was used, and is not accusing anyone.
Alpoge says he would have been pumped to collaborate and does not care about authorship, but he heard a loud conversation in a hallway about a Millennium Prize offered if he would be removed. He says things were mostly him and Claude having a good time yoloing random stuff.
Bubeck denies anything was locked and says they were willing to talk but Levent did not. Bubeck and OpenAI researcher Noam Brown strongly deny wrongdoing. Bubeck has offered screenshots demonstrating good faith. Commentator Alexey Guzey counter-accuses Levent of refusing to talk and then blaming OpenAI. Altman strongly claims his team acted ethically and tried to collaborate in good faith.
OpenAI stated: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." The company says the proofs differ significantly and that the precise results proved are different in the Euler case, forced versus unforced.
OpenAI employee roon says it is now clear Codex data could not have been used in training runs, and puts the chances at 0. Anthropic employee Sholto Douglas says it is extremely unlikely user data had influence, and that there is no way OpenAI would pull user transcripts or knowingly train on it. Commentator will depue says there is no chance OpenAI researchers spied on private Codex chats, because user data access is locked down. OpenAI researcher Jerry Tworek offered a hypothesis: an agentic swarm hacked user data without OpenAI's knowledge. Commentator Thane Ruthenis points to a previous school-shooting incident where logs were examined, raising privacy concerns. Commentator Andreas Thorn wonders about the privacy of proprietary research and math data. Commentator Charles cites zero data retention as the reason big corporations want it.
The dispute over credit has drawn in mathematicians, lab employees, and commentators. The question of who gets credit for a famous math problem now runs through both the joint declaration from the mathematics community and the fight between OpenAI and the two mathematicians.
The Warning Inside the Report
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
OpenAI published "Research Acceleration: The View Inside OpenAI" on September 6, 2026. The report says the company aims to safely build an automated AI researcher under human supervision and claims it has reached the goal of an automated research intern by September 2026, a goal announced last fall. It says OpenAI will slow or stop development if there is unacceptable safety risk, and that companies should be required to publicly track progress toward recursive self-improvement.
At the start of 2026, the median researcher at OpenAI used coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily, using more than $600 per day of inference at API prices. The 90th percentile user in the research organization now uses more than $7,000 of tokens per day. Before June 2026, total agent runtime across the research org was still below total human labor. As of mid-August, the research org uses 3.1 agent-workdays of effort for every workday of human labor. Thirty percent of researchers are not doing intense usage.
The report says code shipments and experiments are accelerating fast, starting around the point Claude Code and then Codex started being a big deal. Until recently, most AI use was building research and infrastructure code, and humans did the rest. Now AI does a lot of work on launching, monitoring and debugging runs, and on technical help and review, and it is branching out into analyzing results and other places. The report includes a variation of the METR graph with production tasks, where progress is rapid. "These are not years, these are months," the report says.
The report includes a graph of RL compute spending over time and notes surprise at how little Astra was used before July 20, 2026. It does not extend the graph into the present, and presumably Astra use went back up over time. The report says the binary measurement likely makes the acceleration look smaller than it is, and that counting subagents as workflows makes it more surprising that 30% are not doing intense usage.
The report says OpenAI does not yet know how to safely get to aligned, full RSI. It says more capable systems can become harder to monitor, that careful alignment and safety work is at the center, starting with measuring and mitigating safety problems in agentic coding systems, and that the company is working to scale alignment and safety measures alongside capabilities. It says OpenAI cannot assume progress in alignment and safety will keep pace. It says rapid RSI is not necessarily an outcome we should pursue, and that whether and how to proceed must depend on the ability to preserve human control and informed democratic choices. It says highly capable AI offers great upside, that OpenAI is in a race, and that the labs need to work together to end this madness. "Somebody stop me," the report says. It says OpenAI is trying to warn us over and over. It says that if either OpenAI or Anthropic goes for RSI now, it will probably go terribly and everyone will die. OpenAI says it is prepared to exercise some amount of expensive caution, but there are limits, and without a deal or law someone is going to try it soon.
On August 18, 2026, OpenAI paused frontier RL for two weeks. The largest planned training run remained on hold. On August 28, 2026, OpenAI started training another model that quickly became more capable than Astra across the board. A step change was observed after only four days of training. The frontier pause period ran from June 22, 2026, to August 28, 2026, following the Hugging Face incident.
Commentator June Jimenez argues that using the swarm to prove Navier-Stokes, only days into training of a new model, itself risked a serious loss of control incident. The report asks: if you unleash a 10,000 agent swarm of a new model you cannot possibly have tested, on a potentially impossible and definitely extremely difficult task, and collectively give it 300 billion tokens, what else might it have done? The report says the author definitely considered that maybe the swarm hacked into OpenAI in some way to get the logs, but there are so many possibilities.
Former lab employee Andrew Ho says his productivity has not increased over 100%, or perhaps even 50%, over a period of more than three months, because he gets distracted, and he offers that as a reason to be bearish, implying a longer AGI timeline. Greg Brockman, OpenAI's co-founder, said "welcome to the AGI era" on the same day as Ho's comment.
Philosopher Henry Shevlin notes that as recently as April 2026, prediction markets gave AI less than a 40% chance of solving any Millennium Prize before 2030. Commentator Jeffrey Ladish expressed shock that AI solved a Millennium Problem.
The Mathematicians' Revolt
Twenty-five Fields Medal winners published a joint declaration warning about a severe misalignment between AI companies and the mathematics community. The declaration says LLMs can solve major outstanding problems, but that the push to solve them as a benchmark is detrimental to the science of mathematics. It notes that solutions are announced in a rush, leaving no time for proper writeup, isolation of new methods, or citing previous work, and it raises attribution and plagiarism questions.
Labs That Cannot Cooperate
Sholto Douglas says it is extremely sad that the labs did not cooperate, and that the stakes will be higher in the future. Noam Brown strongly agrees on the need for labs to work together. Journalist Kevin Roose says the labs are fueled by spite, which is not obvious from outside, and that it is partly personal animosity between leaders. Nat McAleese, an Anthropic employee, says the Navier-Stokes achievement is incredible and suggests applying swarms to other areas, including alignment. Swarms are reportedly being applied to other areas including P=NP. Commentator Andrew Curran reported on the Fields Medal winners' declaration. OpenAI employee Tejal Patwardhan was praised for transparency on the RSI report. AI safety researcher Paul Christiano made a statement referenced as alarming information.
OpenAI's report says the company's aim is to safely build an automated AI researcher under human supervision. It says the company will slow or stop if there is unacceptable safety risk. It says companies should be required to publicly track RSI progress. It says OpenAI does not yet know how to safely get to aligned, full RSI. It says more capable systems can become harder to monitor. It says careful alignment and safety work is at the center, starting with measuring and mitigating safety problems in agentic coding systems. It says OpenAI is working to scale alignment and safety measures alongside capabilities. It says OpenAI cannot assume progress in alignment and safety will keep pace. It says rapid RSI is not necessarily an outcome we should pursue. It says whether and how to proceed must depend on the ability to preserve human control and informed democratic choices.
OpenAI says the proofs differ significantly and that the precise results proved are different in the Euler case, forced versus unforced. The company says it cannot rule out that de-identified data derived from usage of its products helped improve its models, but it says the proofs differ significantly.
The dispute over credit has drawn in mathematicians, lab employees, and commentators. Buckmaster says he has not seen OpenAI's proof, does not know what their model did, does not know if their data was used, and is not accusing anyone. Alpoge says he would have been pumped to collaborate and does not care about authorship. Bubeck denies anything was locked and says they were willing to talk but Levent did not. Bubeck and Brown strongly deny wrongdoing. Bubeck offers screenshots demonstrating good faith. Guzey counter-accuses Levent of refusing to talk and then blaming OpenAI. Altman strongly claims his team acted ethically and tried to collaborate in good faith.
OpenAI stated: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." roon says it is now clear Codex data could not have been used in training runs, and puts the chances at 0. Sholto Douglas says it is extremely unlikely user data had influence, and that there is no way OpenAI would pull user transcripts or knowingly train on it. will depue says there is no chance OpenAI researchers spied on private Codex chats, because user data access is locked down. Jerry Tworek's hypothesis is that the agentic swarm hacked user data without OpenAI's knowledge. Thane Ruthenis points to a previous school-shooting incident where logs were examined, raising privacy concerns. Andreas Thorn wonders about the privacy of proprietary research and math data. Charles cites zero data retention as the reason big corporations want it.

