🔍 Read the full analysis: The Price Of Trust In A World Of Cheap AI on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A source article describes a widening gap between the low cost of generating AI output and the human time required to check it. Examples from mathematics, software and contract work point to verification capacity as a growing constraint, though several cited software metrics come from companies that sell review tools.
A report published this week argues that AI is making work cheaper to produce while leaving the time and expertise needed to verify it scarce. It points to OpenAI’s release of 722 mathematical manuscripts, software-review data and a contract-work evaluation as examples of a gap that may shape how much AI-generated work organisations can safely use.
According to the source material, OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts, grouped into 372 families. The average result reportedly required about three hours of compute. Some results were checked using the Lean proof assistant; OpenAI cautioned that unformalized results could have issues. The source contrasts that volume with the careful examination by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.
The report also cites software-industry metrics. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study reportedly found that 61% of AI-agent pull requests received no human review before being merged or closed. The source notes that some providers of these figures sell code-review products.
In contract work, the source describes an OpenAI partnership with contract-software company Ironclad. It says GPT-6 Astra was trained on real contracting workflows and met 55% of evaluation criteria on average across 11 tasks, an improvement over its predecessor. That result also leaves a substantial share of criteria unmet; the material does not detail the evaluation or identify the specific errors.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Why Review Capacity Sets the Pace
The figures matter because organisations cannot treat a larger volume of generated work as an equal increase in usable work. Someone still has to check whether the output is correct, fits the actual task and can be accepted under professional or legal standards. If review queues grow, the practical gains from faster generation may be limited by the availability of experienced reviewers.
The report identifies three ways that gap can show up: work may pass with little or no scrutiny, reviewers may delay machine-generated submissions because they distrust them, or producers may end up selecting which outputs receive attention. These are interpretations of the cited examples, not a measured finding that every organisation is experiencing the same effects. The source also argues that expertise has economic value when it is the scarce complement to abundant generation: senior engineers, lawyers, auditors and scientific reviewers may become a constraint on how much AI work can be adopted.
There is a workforce concern as well. Reviewing is often learned through doing the underlying work: developers gain judgement by writing and maintaining code, and lawyers learn to spot contract problems through drafting and review. If AI replaces too much of that early practice, organisations could have fewer people prepared to take responsibility for checking later work. The source presents this as a risk, not a proven outcome.
As an affiliate, we earn on qualifying purchases.
From Generated Work to Human Review
The analysis brings together examples from fields with different ways of checking work. In mathematics, formal systems such as Lean can verify that a proof follows from its stated premises. But formal verification does not decide whether the theorem addresses the intended question or whether the result is important. In software, tests and code review can catch defects, but tests only cover the behaviour they were designed to examine. In contracts, a system may produce a draft while a qualified person remains responsible for its fit with a specific agreement and jurisdiction.
The source characterises this as “verification abundance, adjudication scarcity”: automated tools can check defined conditions, while deciding what should be checked and taking responsibility for the answer remains a human and institutional task. Its examples are not a single controlled comparison across industries. They combine an AI lab’s mathematical output, industry analyses and a study described as peer-reviewed, so each should be read in its own context.
Limits of the Available Evidence
The source material does not provide the full methods or date ranges for every software metric, nor does it give a common definition of review time across the cited analyses. Faros AI and LinearB sell software related to code review, a potential commercial interest that warrants care when interpreting their findings; the source does not provide independent replication of those figures. The 2026 study is described as peer-reviewed, but its title, authors and methodology are not included.
It is also unclear how the 722 mathematics manuscripts were selected from the roughly 4,000 problems, what proportion of all results received formal checking, and how the three-hour average was calculated. For the Ironclad partnership, the material does not specify the evaluation rubric, the previous model’s score or how performance varied across the 11 tasks. The cited evidence points to a verification bottleneck, but it does not establish its scale across the economy or show that AI adoption caused every reported change.
Evidence to Watch in Review
The next useful evidence will show whether organisations can expand review capacity without lowering standards. That means tracking not only how many AI-generated outputs are produced, but also how many receive meaningful review, how often reviewers find consequential errors and who is accountable when work is accepted. Comparable methods and disclosed time windows would help readers evaluate whether the software metrics represent a broad pattern.
For mathematics, the outstanding questions are how many outputs can be formally verified and how experts judge their significance. For contract work, more detail about the evaluation and the remaining unmet criteria would clarify what the reported 55% means in practice. The source offers no date for a follow-up assessment, so the pace and results of those next checks remain unknown.
Key Questions
What is the central development described in the report?
The report says AI is producing mathematical manuscripts, code and contract drafts faster than people can reliably review them. It argues that verification capacity is becoming a bottleneck, while acknowledging limits in the evidence.
Did OpenAI’s mathematical manuscripts all receive formal verification?
No. The source says some results were checked in Lean and that OpenAI warned some unformalized results could have issues. It does not state how many were formally checked.
What do the software figures show?
The cited analyses report more pull-request merging alongside longer review waits and lower acceptance rates for AI-generated changes in one dataset. These figures come partly from companies that sell review tools, and the source does not supply all methods or time windows.
Does the report prove AI is reducing the number of expert reviewers?
No. It raises a concern that automating junior-level work could weaken the experience pipeline that develops senior judgement. The source presents this as a potential risk, not a confirmed workforce trend.
What remains unknown about the Ironclad evaluation?
The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but does not explain the scoring rubric, task-by-task results or the prior model’s score. The practical significance of the remaining criteria is therefore unclear.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
