The Price Of Trust In A World Of Cheap AI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Price Of Trust In A World Of Cheap AI on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article describes a widening gap between the low cost of generating AI output and the human time required to check it. Examples from mathematics, software and contract work point to verification capacity as a growing constraint, though several cited software metrics come from companies that sell review tools.

A report published this week argues that AI is making work cheaper to produce while leaving the time and expertise needed to verify it scarce. It points to OpenAI’s release of 722 mathematical manuscripts, software-review data and a contract-work evaluation as examples of a gap that may shape how much AI-generated work organisations can safely use.

According to the source material, OpenAI posed about 4,000 mathematical problems to a model and published 722 manuscripts, grouped into 372 families. The average result reportedly required about three hours of compute. Some results were checked using the Lean proof assistant; OpenAI cautioned that unformalized results could have issues. The source contrasts that volume with the careful examination by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.

The report also cites software-industry metrics. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study reportedly found that 61% of AI-agent pull requests received no human review before being merged or closed. The source notes that some providers of these figures sell code-review products.

In contract work, the source describes an OpenAI partnership with contract-software company Ironclad. It says GPT-6 Astra was trained on real contracting workflows and met 55% of evaluation criteria on average across 11 tasks, an improvement over its predecessor. That result also leaves a substantial share of criteria unmet; the material does not detail the evaluation or identify the specific errors.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA report on AI-generated mathematics, software and contract work argues that output is arriving faster than human verification can keep pace.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Why Review Capacity Sets the Pace

The figures matter because organisations cannot treat a larger volume of generated work as an equal increase in usable work. Someone still has to check whether the output is correct, fits the actual task and can be accepted under professional or legal standards. If review queues grow, the practical gains from faster generation may be limited by the availability of experienced reviewers.

The report identifies three ways that gap can show up: work may pass with little or no scrutiny, reviewers may delay machine-generated submissions because they distrust them, or producers may end up selecting which outputs receive attention. These are interpretations of the cited examples, not a measured finding that every organisation is experiencing the same effects. The source also argues that expertise has economic value when it is the scarce complement to abundant generation: senior engineers, lawyers, auditors and scientific reviewers may become a constraint on how much AI work can be adopted.

There is a workforce concern as well. Reviewing is often learned through doing the underlying work: developers gain judgement by writing and maintaining code, and lawyers learn to spot contract problems through drafting and review. If AI replaces too much of that early practice, organisations could have fewer people prepared to take responsibility for checking later work. The source presents this as a risk, not a proven outcome.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Generated Work to Human Review

The analysis brings together examples from fields with different ways of checking work. In mathematics, formal systems such as Lean can verify that a proof follows from its stated premises. But formal verification does not decide whether the theorem addresses the intended question or whether the result is important. In software, tests and code review can catch defects, but tests only cover the behaviour they were designed to examine. In contracts, a system may produce a draft while a qualified person remains responsible for its fit with a specific agreement and jurisdiction.

The source characterises this as “verification abundance, adjudication scarcity”: automated tools can check defined conditions, while deciding what should be checked and taking responsibility for the answer remains a human and institutional task. Its examples are not a single controlled comparison across industries. They combine an AI lab’s mathematical output, industry analyses and a study described as peer-reviewed, so each should be read in its own context.

Limits of the Available Evidence

The source material does not provide the full methods or date ranges for every software metric, nor does it give a common definition of review time across the cited analyses. Faros AI and LinearB sell software related to code review, a potential commercial interest that warrants care when interpreting their findings; the source does not provide independent replication of those figures. The 2026 study is described as peer-reviewed, but its title, authors and methodology are not included.

It is also unclear how the 722 mathematics manuscripts were selected from the roughly 4,000 problems, what proportion of all results received formal checking, and how the three-hour average was calculated. For the Ironclad partnership, the material does not specify the evaluation rubric, the previous model’s score or how performance varied across the 11 tasks. The cited evidence points to a verification bottleneck, but it does not establish its scale across the economy or show that AI adoption caused every reported change.

Evidence to Watch in Review

The next useful evidence will show whether organisations can expand review capacity without lowering standards. That means tracking not only how many AI-generated outputs are produced, but also how many receive meaningful review, how often reviewers find consequential errors and who is accountable when work is accepted. Comparable methods and disclosed time windows would help readers evaluate whether the software metrics represent a broad pattern.

For mathematics, the outstanding questions are how many outputs can be formally verified and how experts judge their significance. For contract work, more detail about the evaluation and the remaining unmet criteria would clarify what the reported 55% means in practice. The source offers no date for a follow-up assessment, so the pace and results of those next checks remain unknown.

Key Questions

What is the central development described in the report?

The report says AI is producing mathematical manuscripts, code and contract drafts faster than people can reliably review them. It argues that verification capacity is becoming a bottleneck, while acknowledging limits in the evidence.

Did OpenAI’s mathematical manuscripts all receive formal verification?

No. The source says some results were checked in Lean and that OpenAI warned some unformalized results could have issues. It does not state how many were formally checked.

What do the software figures show?

The cited analyses report more pull-request merging alongside longer review waits and lower acceptance rates for AI-generated changes in one dataset. These figures come partly from companies that sell review tools, and the source does not supply all methods or time windows.

Does the report prove AI is reducing the number of expert reviewers?

No. It raises a concern that automating junior-level work could weaken the experience pipeline that develops senior judgement. The source presents this as a potential risk, not a confirmed workforce trend.

What remains unknown about the Ironclad evaluation?

The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but does not explain the scoring rubric, task-by-task results or the prior model’s score. The practical significance of the remaining criteria is therefore unclear.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Security Institute Adds New Executives To Top Team

AI Security Institute has appointed new top executives to strengthen its leadership team, signaling strategic growth in AI security.

Desert Ant Labs: Local, Fast Models That Run On Device

Desert Ant Labs introduces local, fast AI models designed to run directly on devices, signaling a shift toward on-device AI processing amid rising interest.

Kimi K3’s Crucible Run Shows Why AI Agents Need a Real-World Test

Kimi K3 placed second in Firmulate’s company wargame, ahead of three Western models. The results show why AI agents need a test on real business tasks.

What Users Should Know About OpenAI Agents Training Inside Software

OpenAI reports testing an agent on 11 Ironclad workflows. Astra met 55% of rubric criteria on average; its time estimates were simulated.