🔍 Read the full analysis: The Cost Of AI Is Moving From Creation To Quality Control on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
AI systems can generate mathematical manuscripts, software changes and professional-work drafts at increasing volume, but the supplied figures show review remains slower and limited. The data comes from a mix of company research and a peer-reviewed study, and some sources sell review tools, so the precise estimates need care.
AI-generated work is piling up faster than people can verify it, according to a recent analysis drawing on new mathematics and software figures. OpenAI published 722 mathematical manuscripts produced through a research programme this week, while studies of software teams report longer review queues and, in some cases, changes merged without human review. The development matters because organisations may be able to generate more work than they can safely assess.
OpenAI posed about 4,000 mathematical problems to its model and published 722 manuscripts grouped into 372 families, according to the source material. The average result took about three hours of compute to produce. Some manuscripts have been checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could contain issues. Those figures show the scale of production, but do not establish how many results are correct or useful.
The source contrasts that volume with the response to an earlier result from the same programme: a proposed counterexample to an Erdős conjecture that received careful scrutiny from five leading mathematicians. A computer can verify that a proof follows from its stated assumptions, but experts still need to judge whether the claim is correctly framed, significant and informative. The analysis describes this gap as “verification abundance, adjudication scarcity.”
Software data points to a similar pressure. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, which analyzed 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes took 4.6 times longer to begin review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. These measures come from different studies and should not be treated as directly comparable.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Limits AI Output
If those patterns hold across more workplaces, the constraint on using AI may be review time, rather than the ability to produce drafts, code or research. More output is not automatically more usable output: a change that waits for review, a paper whose claims need adjudication or a contract with an unchecked clause still requires qualified human attention.
The analysis also warns of two possible responses to overloaded review teams: work may pass with little scrutiny, or reviewers may delay AI-generated submissions because they expect more problems. LinearB reported that 38% of reviewers deliberately deprioritize AI-generated changes. That figure is a company finding, not a universal measure of reviewer behavior, but it illustrates how distrust can add friction even when an individual change is sound.
Over time, this could shift value toward people who can make reliable judgments and accept responsibility for them: senior engineers, specialist lawyers, auditors and scientific reviewers. That is an interpretation of the reported pattern, not a measured forecast of pay or employment. The analysis raises a further concern: if AI reduces the entry-level work through which people learn to review, the future supply of experienced adjudicators could weaken as demand rises.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Review Gap
The source traces the pattern across mathematics, software and contract work. In each, generation can be partly automated, but checking has limits that go beyond spotting a syntactic error. Formal proof tools can establish that a proof supports a specified theorem; they cannot decide whether the theorem addresses the right question. Likewise, tests check code against selected cases, not whether those cases capture every real need.
The source also cites an OpenAI partnership with contract-software company Ironclad, involving GPT-6 Astra trained on real contracting workflows. On 11 tasks, Astra met an average of 55% of evaluation criteria, according to the supplied material. That indicates progress on the measured tasks, while leaving the remaining criteria for people or other systems to identify. It does not establish how the model performs across all contracts or legal settings.
Company findings require particular care. Faros AI and LinearB sell tools related to software development and review, which may shape how their research is framed. The source says their results point in the same direction, but the figures have different methods, populations and measures. A peer-reviewed study offers a different kind of evidence, though its reported finding also applies to the study sample rather than every software team.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The available figures do not establish a single economy-wide review gap. The studies cover different teams, periods and definitions of review, and the software figures include results from companies that sell related tools. The source does not provide enough methodological detail to reconcile the measures or determine how representative they are.
It is also unclear how many of the 722 mathematical manuscripts are correct, how many will be adopted by researchers, or how often AI-generated work produces serious downstream errors. The Astra evaluation covers 11 tasks, but the supplied material does not describe the criteria, comparison model or performance across other legal workflows. The longer-term effects on hiring, training and reviewer workloads remain uncertain as well.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tracking Review and Training
The next useful evidence will be repeatable measures of review quality and time, not output volume alone. For mathematics, that means reporting which results are formally verified, independently checked and later used. For software and contract work, it means clarifying whether a review occurred, what it caught and what happened after deployment or signing.
Organizations adopting AI will also need to decide how to preserve training in the underlying work. The source argues that junior employees learn judgment through drafting, coding and proving—not only through reviewing machine-generated material. Whether firms change their training practices, and whether AI reduces or expands demand for experienced reviewers, has not yet been established. Those outcomes will show whether higher generation capacity translates into more dependable work.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development?
The analysis says AI is making it easier to generate mathematical manuscripts and software changes, while human verification and review remain slower and limited.
Did OpenAI formally verify all 722 manuscripts?
No. The source says some results were checked in Lean and reports OpenAI’s warning that unformalized results could have issues. It does not say all manuscripts were formally verified.
What did the software studies report?
Faros AI reported more pull requests merged alongside longer review time. LinearB reported longer waits before review began and lower acceptance rates for AI-generated changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests in its sample received no human review before closure or merging.
Are the reported figures conclusive?
No. They come from studies with different methods and populations, and some were produced by companies that sell review-related tools. The figures suggest a possible bottleneck but do not establish a universal rate.
What remains unknown?
It is not yet clear how representative the review findings are, how many AI-generated results contain consequential errors, or how workplaces will adapt training and staffing as AI use grows.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
