🔍 Read the full analysis: More AI Production Means More Work For The Referees on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A report from ThorstenMeyerAI.com argues that AI is making it faster to produce mathematical research, code and professional drafts, while human verification remains time-consuming and limited. It cites examples from OpenAI’s mathematics work and software-review datasets, but some figures come from vendors and their scope and methods require care.
A report published this week argues that AI is increasing the volume of work produced faster than people can verify it, pointing to OpenAI’s publication of 722 mathematical manuscripts and software-industry data showing longer review waits for AI-generated code. The report’s central concern is that human reviewers may become a constraint on how much AI-produced work organisations can safely use.
The source says OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts, grouped into 372 families. It reports that producing an average result took about three hours of compute. Some results were formally checked using Lean, a proof-assistant system; OpenAI cautioned that some results not formalised in that system could have issues. The figures describe output, not how many results have been independently accepted as significant or correct.
The report contrasts that volume with the response to an earlier result from the same programme: a proposed counterexample to an Erdős conjecture was carefully checked by five leading mathematicians. That comparison illustrates the source’s distinction between generating candidate results and establishing that they are valid, relevant and properly interpreted. The report characterises this as “verification abundance, adjudication scarcity”.
Software figures cited by the source point in a similar direction, though they come from separate studies with different methods. Faros AI reported that teams in high-AI-adoption periods merged 98% more pull requests while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that several providers of the data sell code-review products, which is a reason to interpret their measurements with care.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If the pattern holds, the main limit on AI adoption may not be the ability to generate more material but the ability to check it well. A team can produce extra code, research or contract drafts, but each item still has to be assessed against the right requirements, checked for errors and approved by someone with appropriate responsibility. That review takes time and specialist knowledge.
The source also describes risks when review capacity falls short: work can pass with little scrutiny, reviewers can deprioritise AI-generated submissions as a group, or the producer’s own selection process can substitute for independent review. These are possible organisational responses discussed in the report, not outcomes established for every workplace. The consequences could matter in fields where errors affect users, clients, research conclusions or legal obligations.
The report’s longer-term concern is the supply of experienced reviewers. If junior workers spend less time drafting code, contracts or proofs themselves, they may have fewer opportunities to learn how to spot errors in that work. That creates a potential training problem: organisations could need more judgement at the same time that fewer people are gaining the experience that builds it.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Bottleneck
The report draws its examples from mathematics, software development and contract work. In mathematics, formal proof tools can check whether a proof establishes a stated theorem, but they do not decide by themselves whether the theorem is the intended one or whether the result matters. For software, automated tests check the cases they cover; they cannot establish that those tests capture every requirement. In contract work, a draft may meet many evaluation criteria and still require a professional to catch a missed approval rule or an unsuitable clause.
As another example, the source describes OpenAI’s partnership with contract-software company Ironclad and an evaluation of a model it calls GPT-6 Astra. It says the model met an average of 55% of evaluation criteria across 11 tasks, an improvement over its predecessor. That is a reported task-evaluation result, not evidence that the model can independently complete legal work. The source says the remaining criteria still require attention before the drafts can be relied on.
Across the examples, the source’s argument is that verification is not simply a second generation task. Reviewers must judge whether the output answers the right question and whether it is safe and appropriate to use. The report also says responsibility remains with people and institutions, such as engineers signing off designs, lawyers handling contracts and authors answering for research.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The source does not provide enough detail to determine whether the cited software metrics are directly comparable, how teams and review times were defined, or whether the reported differences are caused by AI adoption. It notes that several datasets come from companies selling review tools; the figures should be treated as attributed findings rather than a universal measure of software quality.
For the mathematics results, the source does not give a full breakdown of how many manuscripts were independently verified, rejected or later revised. It also does not establish that three hours of compute represents the full human effort involved in producing or checking each result. The account of the earlier conjecture counterexample does not, on its own, establish the status of every manuscript in the larger collection.
The source’s example of 55% of evaluation criteria across 11 contract tasks is also limited to that stated evaluation. It does not say how the criteria were weighted or show how the model performs across the broader range of legal work. The scale of any future shortage of qualified reviewers remains uncertain.
mathematical proof assistant software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tracking Output and Oversight
The next useful evidence will show whether review times, rejection rates and unreviewed merges continue to change as more teams adopt AI tools. Independent studies with clear definitions and comparable time periods could help establish whether the vendor-reported software patterns apply across industries, rather than only within the datasets cited here.
For AI-generated mathematics and contract work, further reporting on independent verification, errors found after release and human approval would clarify how much of the apparent productivity gain is usable in practice. The report offers no announced policy change or fixed timeline for resolving the reviewer-capacity issue. For now, the key question is whether organisations can expand production while maintaining enough skilled, accountable human oversight.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described in the report?
The report argues that AI can generate work faster than people can verify it, citing OpenAI’s 722 mathematical manuscripts and software-review data.
Does the mathematics figure mean all 722 manuscripts were verified?
No. The source says some results were formally checked in Lean and quotes OpenAI warning that some unformalized results could have issues. It does not give a complete independent-verification count.
What did the software data show?
The source cites separate analyses reporting longer waits for review of AI-generated changes and, in one dataset, lower acceptance rates than for human-written changes. Those findings have different methods and should not be treated as one universal measure.
Why might human review remain necessary if AI can check AI?
Automated tools can check specified proofs, tests or evaluation criteria, but people may still need to decide whether the specification is correct, whether important cases are missing and who is accountable for the result.
What remains unknown?
The available material does not establish how broadly the cited metrics apply, how many mathematical manuscripts have been independently verified, or whether the potential shortage of experienced reviewers is already affecting organisations at scale.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
