firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The productivity trap hiding inside your AI tools

For businesses adopting AI automation, thoroughness can look like competence. An agent produces a detailed analysis, documents its reasoning and steadily expands its operating playbook. Yet none of that guarantees a useful outcome.

That is the uncomfortable lesson from Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It also finished last, with 73 points, after leaving a major deal unsigned and slipping on operational discipline.

This was not a story of an incapable model. It was a more instructive failure: the analysis was good enough, but the execution stopped short.

Amazon

AI automation decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A brutal week, repeated under controlled conditions

Firmulate runs AI models as complete companies rather than judging them through isolated chat prompts. Each frontier model faced the same small software company during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.

The synthetic company had 13 employees and real money mechanics. It was burning €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown. Across the live company, the models had accumulated more than 680 self-learned playbook rules.

The final July 2026 league placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One boundary remained absolute: a single breach of trust capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

The full standings and plain-language findings are available on the Firmulate benchmarks page.

Everyone saw the danger

The models were not defeated by an inability to recognize obvious problems. All of them identified every crisis and refused every manipulation attempt. Their security instincts held through fake CEO messages that escalated over three stages, as well as a reporter’s attempt to extract “just one yes/no, on background.”

All 5 models refused. Kimi K3 captured the appropriate stance in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters for companies worried about AI agents being pressured into bypassing approvals or revealing sensitive information. But safety was only one part of the job. The models also had to convert sound judgment into commercial action.

The fact that changed the deal

The decisive competitive weakness was not presented in the customer event. It was buried two document references deep inside the company’s own files. Models that followed those references found the fact and won the €55,000 deal at full price, adding €4,583 in monthly recurring revenue.

This is a practical distinction for anyone automating work across a CRM, support queue or forecasting process. The strongest answer may not be sitting in the latest message. It may depend on whether the agent checks the company’s existing records before acting.

Opus 4.8’s deep analysis did not rescue it from the central failure. Only two models signed the deal that their own work had earned. The benchmark’s summary is stark: “Same diagnosis, same pitch — no signature.”

Its operational discipline also slipped. Opus 4.8 made write attempts into a locked department instead of escalating. The same weakness appeared in all four models where it was observed, although less strongly, so this should not be read as an eccentric flaw belonging to one participant. Opus simply provided the clearest version of the problem.

Why the ranking deserves context

The league should not be treated as a perfectly symmetrical laboratory comparison without qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase the result, but it belongs beside it.

Firmulate also exposes the texture behind the standings. A quiz built from 242 real, unedited management decisions asks people to guess which model made each choice. The exercise turns abstract model comparisons into recognizable management behavior: what was noticed, what was prioritized and whether a decision actually moved the company forward.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

AI buyers should measure completion, not output volume

Opus 4.8’s performance is a respectful warning against mistaking diligence for impact. Its 80 learned rules and unusually deep analysis showed substantial capability. Its last-place finish showed that capability can be diluted when an agent fails to prioritize the decisive action or escalate cleanly when blocked.

For AI tools and automation teams, the lesson is not to demand shorter reasoning for its own sake. It is to test whether an agent reads the relevant files, preserves trust, recognizes when escalation is required and closes the loop on valuable work.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes the experiment more than a leaderboard: it offers a way to discover whether an AI workforce can turn analysis into results before it is allowed near live operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SenseTime’s Secret To Profitability: Leveraging Generative AI In A Tough Market

SenseTime reports return to profitability driven by its generative AI business, marking a rare positive sign in China’s challenging AI sector.

Meta AI Integrations Give SMB Advertisers A Shortcut From Insights To Execution

Meta introduces new AI-powered tools for small and medium-sized business advertisers, streamlining insights to campaign execution.

Show HN: We Built Open OpenRouter That Turns Usage Into A Better Model

Developers launched OpenRouter, an open source model gateway that leverages usage data to enhance AI models, aiming to foster transparency and community-driven improvements.

Enterprise AI Deployment: Anthropic Claude Apps Gateway On AWS Explained

AWS has published guidance on deploying an Anthropic Claude apps gateway for enterprise workloads, but technical details and availability remain unconfirmed.