The Cost Of AI Is Moving From Creation To Quality Control
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Cost Of AI Is Moving From Creation To Quality Control on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI systems can generate mathematical manuscripts, software changes and professional-work drafts at increasing volume, but the supplied figures show review remains slower and limited. The data comes from a mix of company research and a peer-reviewed study, and some sources sell review tools, so the precise estimates need care.

AI-generated work is piling up faster than people can verify it, according to a recent analysis drawing on new mathematics and software figures. OpenAI published 722 mathematical manuscripts produced through a research programme this week, while studies of software teams report longer review queues and, in some cases, changes merged without human review. The development matters because organisations may be able to generate more work than they can safely assess.

OpenAI posed about 4,000 mathematical problems to its model and published 722 manuscripts grouped into 372 families, according to the source material. The average result took about three hours of compute to produce. Some manuscripts have been checked using Lean, a proof-assistant system; OpenAI cautioned that unformalized results could contain issues. Those figures show the scale of production, but do not establish how many results are correct or useful.

The source contrasts that volume with the response to an earlier result from the same programme: a proposed counterexample to an Erdős conjecture that received careful scrutiny from five leading mathematicians. A computer can verify that a proof follows from its stated assumptions, but experts still need to judge whether the claim is correctly framed, significant and informative. The analysis describes this gap as “verification abundance, adjudication scarcity.”

Software data points to a similar pressure. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, which analyzed 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes took 4.6 times longer to begin review and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. These measures come from different studies and should not be treated as directly comparable.

At a glance
analysisWhen: Published this week; software and resea…
The developmentA new analysis of AI-generated mathematics, software and contract work argues that the cost and capacity bottleneck is shifting from producing output to verifying it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Limits AI Output

If those patterns hold across more workplaces, the constraint on using AI may be review time, rather than the ability to produce drafts, code or research. More output is not automatically more usable output: a change that waits for review, a paper whose claims need adjudication or a contract with an unchecked clause still requires qualified human attention.

The analysis also warns of two possible responses to overloaded review teams: work may pass with little scrutiny, or reviewers may delay AI-generated submissions because they expect more problems. LinearB reported that 38% of reviewers deliberately deprioritize AI-generated changes. That figure is a company finding, not a universal measure of reviewer behavior, but it illustrates how distrust can add friction even when an individual change is sound.

Over time, this could shift value toward people who can make reliable judgments and accept responsibility for them: senior engineers, specialist lawyers, auditors and scientific reviewers. That is an interpretation of the reported pattern, not a measured forecast of pay or employment. The analysis raises a further concern: if AI reduces the entry-level work through which people learn to review, the future supply of experienced adjudicators could weaken as demand rises.

Amazon

AI review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Review Gap

The source traces the pattern across mathematics, software and contract work. In each, generation can be partly automated, but checking has limits that go beyond spotting a syntactic error. Formal proof tools can establish that a proof supports a specified theorem; they cannot decide whether the theorem addresses the right question. Likewise, tests check code against selected cases, not whether those cases capture every real need.

The source also cites an OpenAI partnership with contract-software company Ironclad, involving GPT-6 Astra trained on real contracting workflows. On 11 tasks, Astra met an average of 55% of evaluation criteria, according to the supplied material. That indicates progress on the measured tasks, while leaving the remaining criteria for people or other systems to identify. It does not establish how the model performs across all contracts or legal settings.

Company findings require particular care. Faros AI and LinearB sell tools related to software development and review, which may shape how their research is framed. The source says their results point in the same direction, but the figures have different methods, populations and measures. A peer-reviewed study offers a different kind of evidence, though its reported finding also applies to the study sample rather than every software team.

Amazon

code review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Enough?

The available figures do not establish a single economy-wide review gap. The studies cover different teams, periods and definitions of review, and the software figures include results from companies that sell related tools. The source does not provide enough methodological detail to reconcile the measures or determine how representative they are.

It is also unclear how many of the 722 mathematical manuscripts are correct, how many will be adopted by researchers, or how often AI-generated work produces serious downstream errors. The Astra evaluation covers 11 tasks, but the supplied material does not describe the criteria, comparison model or performance across other legal workflows. The longer-term effects on hiring, training and reviewer workloads remain uncertain as well.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The next useful evidence will be repeatable measures of review quality and time, not output volume alone. For mathematics, that means reporting which results are formally verified, independently checked and later used. For software and contract work, it means clarifying whether a review occurred, what it caught and what happened after deployment or signing.

Organizations adopting AI will also need to decide how to preserve training in the underlying work. The source argues that junior employees learn judgment through drafting, coding and proving—not only through reviewing machine-generated material. Whether firms change their training practices, and whether AI reduces or expands demand for experienced reviewers, has not yet been established. Those outcomes will show whether higher generation capacity translates into more dependable work.

Amazon

AI quality control tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development?

The analysis says AI is making it easier to generate mathematical manuscripts and software changes, while human verification and review remain slower and limited.

Did OpenAI formally verify all 722 manuscripts?

No. The source says some results were checked in Lean and reports OpenAI’s warning that unformalized results could have issues. It does not say all manuscripts were formally verified.

What did the software studies report?

Faros AI reported more pull requests merged alongside longer review time. LinearB reported longer waits before review began and lower acceptance rates for AI-generated changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests in its sample received no human review before closure or merging.

Are the reported figures conclusive?

No. They come from studies with different methods and populations, and some were produced by companies that sell review-related tools. The figures suggest a possible bottleneck but do not establish a universal rate.

What remains unknown?

It is not yet clear how representative the review findings are, how many AI-generated results contain consequential errors, or how workplaces will adapt training and staffing as AI use grows.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A September 2026 Look At My AI Tools And Workflow

A developer’s Sept 2026 AI stack: Opus 5.5 builds, GPT-6.1 Sol reviews at $0.39 a task, as six top models converge within 20 index points.

AI-Powered “Robot Scientist” Makes a Chemistry Breakthrough on Its Own

From autonomous experimentation to groundbreaking discoveries, this AI robot scientist’s breakthrough in chemistry challenges our understanding of scientific progress and ethics.

How Robots Learn From Trial and Error

AIThis post was created with the assistance of artificial intelligence (AI).You see,…

How Recommendation Algorithms Quietly Shape What You See

What you see online is subtly crafted by algorithms that influence your choices—discover how to recognize and challenge these unseen forces.