The live catalogue is 719 math manuscripts. Three others are already gone over a sign error.

2026-10-11

On 6 October 2026 OpenAI put a pile of mathematics on GitHub (a public file host that keeps every change, used more often for code than for papers) and called it progress from an internal frontier model (a model the company has not released, stronger than the one in the public chat product). The live catalogue, in the repository’s own readme, is 719 manuscripts in 372 families, and about 42 percent of the top-line results have a proof a computer accepted. Three other manuscripts are already withdrawn. The reason is a sign error (a plus where the argument needed a minus, or the reverse). A title is not a result. This is how to tell which sentence has a checker.

Start one step before the lab words. A claim (a sentence that says something is true) is not yet a proof (a chain of steps from definitions you already accepted to that sentence). A proof is not yet a preprint (the write-up, posted so other people can read it before a journal has accepted it). A journal adds referees (people the journal asks to try to break the argument before it will print it). None of those words means “a model said so.” A model can type all four and still be wrong at step two.

From claim to journal paper: a claim becomes a proof, the proof becomes a preprint posted before a journal accepts it, and a journal paper is one referees have tried to break

Why did OpenAI post its math preprints on GitHub, not arXiv?

Simple mindmap of the OpenAI math preprints: 719 manuscripts, Lean checks statements, three withdrawn, arXiv two-a-month cap, famous names are special cases, check history.md

Mathematicians usually post that write-up on arXiv (a free server, run for researchers, that shows new preprints after volunteer moderators look at them). On 1 October 2026 arXiv capped every submitter at two papers per calendar month, and at three waiting at once. Rejected papers still count, because a rejection still costs a moderator’s time. arXiv’s own figures: 9,869 submissions in September 2016, 20,569 in September 2024, 40,363 in September 2026, and almost 9,000 support tickets that last month. Its blog says AI tools are making it easy to flood repositories with low-value papers. At two a month, one author would need about 30 years to post 719 manuscripts. GitHub was the door that could open in a day. That door has no referee standing in it.

Why the OpenAI math preprints went to GitHub: arXiv caps submitters at 2 a month, 719 manuscripts would take about 30 years, GitHub took the pile in a day with no moderator

What is in the OpenAI math repository?

Here is what OpenAI itself says is in the room. The readme at github.com/openai/math, under an Apache-2.0 license (a standard permission slip that lets other people copy the files, with conditions), says the catalogue is 719 manuscripts in 372 families. A family (a bundle of related papers around one result: the main argument, side arguments, consequences, or a second proof) is the unit to count if you care about results rather than PDF files. The same readme says the model was posed approximately 4,000 problems, and that a typical result used three hours of ChatGPT Pro thinking compute on that unreleased model. Ten families also have an abridged reasoning summary (a shortened write-up of how the model got there, not the full scratch work). Two results did not follow the usual recipe: a zero-free region for the Riemann zeta function, whose write-up was edited by a person for readability, and a claimed proof of the Hodge conjecture for CM abelian varieties.

What the repo saysThe plain reading
719 manuscripts, 372 familiesMore PDFs than results. Count families if you are counting claims.
About 4,000 problems posedMost attempts did not become a manuscript. The pile is the subset they chose to publish.
Three hours of ChatGPT Pro thinking compute, on averageA cost comparison. It is not a button in the public product. The model is unreleased.
~42 percent of top-line results formalized, stated as 300 / 719 in the history fileA computer has accepted a formal proof for about two in five headline results. The rest are English.
“Some of the unformalized results could have issues.”OpenAI’s own warning. It is not a critic’s insult.

Did OpenAI prove the Riemann hypothesis or the Hodge conjecture?

Those two famous names need a smaller box before anyone screenshots them. The Riemann hypothesis (the claim that the zeros of a particular function, the zeta function, all sit on one vertical line, which would pin down how the prime numbers are spaced) is not what the readme says they proved. A zero-free region (a proof that no zeros sit in some strip, short of the full line) is a smaller claim. The one they flag was human-edited, and the strip they name is the real part of s greater than 11/12. The Hodge conjecture (a claim about which shapes inside a geometric space come from polynomial equations) is also not one sentence. “For CM abelian varieties” is one family of spaces. A proof there, if it holds, is not a proof for every space. Nature reported that this batch did not include solutions of the five remaining Millennium Prize Problems (seven famous problems with a cash prize; one was solved years ago).

Famous problem names in OpenAI math preprints are special cases: a zeta zero-free region is not the Riemann hypothesis, CM abelian varieties is not every space for Hodge

What does Lean formalization actually check?

Now the word that decides whether you can trust a file you cannot follow. You already know a compiler (a program that reads code and refuses to finish if the code breaks the language’s rules). Lean (a language in which a proof is code, so that same kind of refusal applies to a mathematical step) is a compiler for arguments. Formalization (rewriting the argument in Lean so the checker, not a tired reader, decides whether each step follows) is what the “42 percent” is counting. The history file states the fraction as 300 / 719. The checker accepts the Lean statement you typed. It does not check that the English title means the same thing as that statement. A perfect Lean file can still sit under a boastful PDF. An unformalized PDF can still be right. OpenAI’s sentence covers the second risk and not the first: “Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly.”

How to trust an OpenAI math preprint: with no Lean proof treat it as a lead; with Lean, check the formal statement matches the English claim before calling it machine-checked

Why were three OpenAI math preprints withdrawn?

The warning was not boilerplate. The history file says a sign error in “Algebraicity of Weil classes on split abelian eightfolds” breaks a cancellation step, and that the same construction was used by two other papers. OpenAI withdrew all three: that paper, “Algebraicity of Kuga-Satake Correspondences for K3 Surfaces,” and “The rational Hodge conjecture for products of K3 surfaces.” Withdrawn does not mean “the Hodge conjecture is false.” It means those three write-ups do not prove what they offered. The file also says 14 other manuscripts were revised (proof repairs, corrected statements, clearer hypotheses, one obsolete citation), and 13 more were updated so they cite the revised companions. Early descriptions of the drop, including OpenAI’s own developer forum on 6 October, said 722 manuscripts. The live readme says 719. Three left. The archive still holds them, with a notice. That is the right way to retract. It is also evidence that a clean-looking paper in this pile can be dead within days.

A sign error is not a matter of taste. Two quantities that were supposed to cancel do not cancel if one of them has the wrong sign. Everything that used the cancellation falls with it. That is why two other papers left with the first one.

One sign error in a cancellation withdrew the Weil classes, Kuga-Satake and rational Hodge for K3 products papers, while 14 other manuscripts were revised and 13 citations updated

What are mathematicians saying about the drop?

Both reactions you will see online can be honest. Kevin Buzzard, a mathematician at Imperial College London, told The Verge he either has to read possibly-not-correct slop, wait for someone else to, or wait for a formalization before he can say the result is even correct. Ursula Martin, at Oxford, told Nature the release was tossing the community a messy first draft and expecting them to clean it up. Andrew Sutherland, at MIT, told Scientific American he expects most of the proofs to be correct or correctable, and also expects some mistakes, possibly serious ones. Peter Woit, at Columbia, called the drop a slopocalypse and wrote that the pile was so badly organized he had been scrolling through a fraction of the 722 and thought he was looking at the whole. Other mathematicians, quoted by Nature and The Verge, called it a historic moment. A checked proof of a real problem is allowed to be important. A release that hands hundreds of half-checked PDFs to a field that reads slowly is allowed to be a mess. Those are not opposite facts.

Did OpenAI follow the AGMAI release rules?

The advisory group OpenAI cites asked for something the drop does not do. On 29 September 2026 the Advisory Group on Mathematics and Artificial Intelligence (AGMAI, researchers convened at the Institute for Advanced Study) published release rules. The opening practical point is that they do not endorse labs testing advanced problems on proprietary models (models outsiders cannot run), and they ask labs to stop. OpenAI’s 6 October post says it drew on that advice, and in the same post says it is important to keep evaluating internal frontier models on open research problems. TechCrunch reported that AGMAI’s follow-up was that the community, not the group, would judge whether the recommendations were followed. Consulting a committee is not the same as doing the thing the committee put first.

AGMAI asked labs on 29 September to stop testing hard problems on models outsiders cannot run; OpenAI said it drew on that advice yet will keep evaluating internal models

How should you read one of the OpenAI math preprints?

If you are going to open one file, use this order. It is shorter than reading the pile, and it refuses the two mistakes the pile invites: believing a title, and ignoring a checker.

  1. Open the manuscript map, not a screenshot. Find the family. A family can hold a main paper plus companions. Quoting the companion as if it were the theorem is how a small lemma (a helper result, proved so a larger result can use it) becomes a press headline.
  2. Read history.md before you cite. Three titles already point at a gap. Fourteen more have been patched. A link to the first commit is a link to a version OpenAI may already have replaced.
  3. If there is no Lean file, stop calling it a theorem in your notes. Call it a lead. That is the word that matches OpenAI’s own warning.
  4. If there is a Lean file, write down the formal statement in one sentence of your own, then look back at the title. The checker signed the statement. It did not sign the title. This gap (the formal statement and the English claim coming apart) is the bug a compiler cannot see, because both sides are not in the file it was given.
  5. If the title contains a famous problem, look for the qualifier. “A zero-free region” is not “the Riemann hypothesis.” “For CM abelian varieties” is not “the Hodge conjecture.” “For products of K3 surfaces” was a third, narrower claim, and that manuscript is withdrawn.
  6. Do not paste an unformalized PDF into the next training set as if it were a textbook. The next model will quote the sign error and attach a citation. A citation feels like a check. It is a pointer.
  7. The ten reasoning summaries are the readable on-ramp. They are abridged. They are not a substitute for the Lean file, and they are not the full scratch work.
Reading order for one OpenAI math preprint: find the family, check history.md, check for a Lean file, compare the formal statement to the title, cite the live version

What the OpenAI math release is not

What not to take out of this week. This is not “mathematics is finished,” and it is not “the folder is fake.” About two in five headline results come with a proof a computer accepted, the company is revising in public, and the old versions stay up. It is also not a model you can run. You can read the PDFs. You cannot ask the system that wrote them to explain a line. The three-hour figure is a compute comparison against ChatGPT Pro thinking, measured on the internal model. It is not a recipe. And a special-case theorem with a famous surname is not the prize problem. Nature’s line is the one to keep next to any viral claim: this batch did not solve the five remaining Millennium Prize Problems.

The sentence worth keeping: a computer check applies to a formal statement, a history file applies to a changing folder, and a title applies to neither until a person matches them.

Common questions about the OpenAI math preprints

How many OpenAI math preprints are there?

The live readme lists 719 manuscripts in 372 families, drawn from about 4,000 problems posed to an unreleased internal model. Early descriptions, including OpenAI’s developer forum on 6 October, said 722.

How many of the OpenAI math results are checked in Lean?

About 42 percent of top-line results, stated as 300 / 719 in the history file. The rest are English write-ups, and OpenAI warns some unformalized results could have issues.

Why were three OpenAI math papers withdrawn?

A sign error in “Algebraicity of Weil classes on split abelian eightfolds” breaks a cancellation step, and two other papers used the same construction. All three were withdrawn, and 14 other manuscripts were revised.

Did OpenAI solve the Riemann hypothesis?

No. The readme claims a zero-free region for the zeta function, a smaller result, and that write-up was edited by a person. Nature reported the batch did not solve the remaining Millennium Prize Problems.

Why did OpenAI use GitHub instead of arXiv?

Since 1 October 2026 arXiv caps each submitter at two papers a month and three waiting at once. At that rate, 719 manuscripts would take one author about 30 years.

Leave a comment