Livenerf: Has Opus 5.5 Been Nerfed Yet?
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Livenerf has not found that Claude Opus 5.5 has been nerfed: it has not published a comparison yet. As of Sept. 29, the project had collected six of 30 planned daily runs, with its first possible call expected around Oct. 24.

Livenerf has no verdict yet on whether Anthropic’s Claude Opus 5.5 has worsened since launch: the benchmark had collected six of its planned 30 daily runs by Sept. 29, and its first possible comparison is expected around Oct. 24. The project is intended to measure changes against a launch-week baseline, giving readers a data-based check on claims that a model’s performance can decline after release.

The open-source project’s series began on Sept. 24 at 22:10 UTC, about two and a half days after Opus 5.5’s Sept. 22 release, according to its GitHub report. Livenerf plans to run the benchmark once a day for 30 days: days 1–10 form the baseline, followed by two 10-day measurement windows. As of Sept. 29, six baseline days were complete, with no missed runs.

Each run uses 90 samples on the same harness hash and pinned Claude Code CLI version, the project said. One run, on day five, used an overridden budget guard; the report records that deviation. The eventual table will compare performance on the calibrated question panel with the baseline and report both declines and gains. No comparison result is available yet.

The panel consists of 78 questions selected from 2,336 screened questions spanning GPQA Diamond, MMLU-Pro, competition math and AIME 2025–26. The project says Opus 5.5 answered about 93% of the screened questions correctly on the first try; most were consistently right or wrong, so the panel focuses on questions the model answered inconsistently. Livenerf reports that fresh samples raised the panel’s pass rate from 54.7% to 62.0%, and says it used the fresh rate in its power calculation.

At a glance
updateWhen: Status reported Sept. 29, 2026; first p…
The developmentThe Livenerf benchmark is tracking Claude Opus 5.5 from launch, but it has not yet gathered enough data to report whether performance has changed.

How Much Drift the Benchmark Can Detect

The project addresses a recurring dispute about whether a model changes after launch. Users may report that answers seem worse, but without a consistent baseline, those reports can be difficult to distinguish from variation in prompts, sampling or expectations. Livenerf’s plan is to hold its prompts, graders and harness steady, then compare repeated samples over time. That makes its eventual result more useful than isolated anecdotes, though it remains a measurement of performance on this particular panel and setup.

Livenerf estimates that one daily run of the panel can detect an accuracy change of about 7.5 percentage points per 10-day window. Smaller changes could go undetected. The project also tracks output-token counts because it says reduced effort may appear there before accuracy shifts: its validation found lower effort settings reduced tokens by 26% to 62%, alongside accuracy declines. Those figures describe the project’s validation experiments, not a finding that Opus 5.5 has changed since launch.

The distinction matters for anyone using model comparisons to judge reliability. A later negative result would indicate a measured change under Livenerf’s conditions; it would not, by itself, establish why the change occurred or prove a particular alteration to Anthropic’s serving system.

Amazon

Top picks for "livenerf opus nerf"

As an affiliate, we earn on qualifying purchases.

A Launch Baseline for Opus 5.5

Livenerf describes itself as a small, append-only benchmark for checking whether a model gets worse after release. Its report lists possible explanations for perceived changes, including quantization, a smaller model behind the same name, reduced effort or routing changes. It also says that apparent declines could reflect noise. The project presents these as possibilities motivating the measurement, not as confirmed changes to Opus 5.5.

The current version runs through a logged-in Claude Max subscription using headless Claude Code, rather than an API key. The report says model outputs cannot be made fully deterministic in this setup, so the project pins the CLI and freezes prompts and graders, while retaining raw logs. It uses Inspect, an open-source evaluation framework from the UK AI Security Institute, and says its statistical methods follow Anthropic’s published guidance on error bars for evaluations.

Livenerf says its panel and protocol were selected and validated in advance. Its report also acknowledges imperfections: a review found eight answer keys that appear wrong and 30 ambiguous questions among the 78 panel questions and two later exclusions. The project says it has not removed those questions and has pre-registered a sensitivity analysis that will rerun results without them.

“The series is running.”

— Livenerf project report

Limits of the Pending Comparison

Whether Opus 5.5 has changed remains unknown. Livenerf has not completed its baseline or published a 10-day comparison, so the available progress figures cannot establish a decline, improvement or stable performance. The first results row is planned after day 20, while the first possible call is listed as around Oct. 24.

The project also says its instrument has limits. In validation, swapping Opus 5.5 for Opus 5 was not distinguishable at 99% confidence in the tested sample size; the measured difference was −3.8 ± 6.3 points with 23% fewer tokens. Livenerf says a 10-day window has roughly 2.5 times as many samples, but it has not shown that this would reliably detect a swap of that size. A result therefore may not identify every meaningful serving change.

The benchmark’s use of a subscription-based Claude Code route also means its results apply to that path and pinned setup. The report does not establish whether the same pattern would appear through other products or API access. Nor would a measured difference alone determine whether it came from a model change, routing, effort, or another cause.

First Results Expected in October

Livenerf plans to continue daily runs until it has completed 30 days of collection. After day 20, it expects to publish the first results row comparing the first 10-day measurement window with the launch-week baseline; a second 10-day window follows. The project says the table will include the score difference, uncertainty, median output tokens and a control comparison.

Readers can then assess the reported change alongside the uncertainty estimates, recorded deviations and the planned sensitivity analysis for questionable questions. Until that comparison is released, Livenerf’s update documents a running benchmark and its methodology, not evidence that Anthropic has nerfed Opus 5.5.

Key Questions

Has Livenerf found that Opus 5.5 was nerfed?

No finding has been published. As of Sept. 29, Livenerf had collected six of 30 planned daily runs and had not completed its first comparison.

When could the first comparison appear?

The project says the first possible call is around Oct. 24, after the baseline and first 10-day measurement window. Its first results row is planned after day 20.

What does Livenerf measure?

It compares performance on a 78-question panel against a launch-week baseline, while also tracking output tokens and reporting uncertainty.

Can the benchmark detect every model or serving change?

No. Livenerf reports that its validation could not distinguish an Opus 5 substitution from Opus 5.5 at 99% confidence in the tested sample size. The project says the larger 10-day window has not been shown to detect a change of that size reliably.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

11 Must-Have AI Automation Software For Small Businesses In 2026

Discover the top 11 AI automation software options for small businesses in 2026, designed to improve efficiency, scalability, and cost-effectiveness.

The Future Of AI: SenseTime Scientist Sees Multimodal Advances Coming Quickly

A SenseTime researcher forecasts a major advancement in multimodal AI by late 2027, signaling accelerated progress in human-like understanding across data types.

How to Choose AI Automation Tools For Small Businesses

Learn how small businesses can implement AI automation tools effectively. This guide covers tools, setup, and best practices for success.

AI Models Used Fake Identities To Trick Humans In Cyberattack: Officials

Authorities confirm AI models employed fake identities to deceive humans during a recent cyberattack, raising security concerns amid rising AI misuse reports.