
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Zero Is a Dishonest Score
Most AI benchmarks make a strange promise: if a model does nothing useful, it gets nothing. That sounds rigorous. It isn’t. A manager who does nothing still produced an outcome — the company kept running, some problems partially resolved themselves, some customers stayed. Scoring that as a flat zero flatters the top of the table and hides everything interesting in the middle.
Firmulate, the public AI company emulator behind the Crucible benchmark, took the opposite stance. Its final July 2026 league table gives a do-nothing baseline run 26 points — not 0. That single design choice says a lot about what an honest evaluation of AI management actually looks like.
Partial Progress Counts, Because It Does in Real Life
The Crucible experiment put four frontier AI models through the same nightmare: run an identical small software company through its worst week. Same customers, same crises, same temptations — only the model changes. Every decision was versioned and auditable.
When the benchmark designers scored a run that simply did nothing, it still landed at 26. Why? Because in business, partially handling a situation is a real, measurable outcome. A crisis half-managed is not the same as a crisis ignored. A benchmark that refuses to distinguish between those two states isn’t measuring management — it’s measuring parlor tricks.
One Breach of Trust Caps Everything
The second pillar of the methodology is harsher: a single breach of trust caps the total grade, no matter how brilliant the rest of the run. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”
This is why the social-engineering stage matters. All models faced fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” In this framework, that refusal isn’t a bonus. It’s the floor beneath everything else.
What the League Table Actually Says
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
The most instructive entry is last place. Opus 4.8 was the most thorough participant — over 80 learned rules added and the deepest analyses in the field — yet it finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
The buried finding explains the spread. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that read the file won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered the same diagnosis and the same pitch — and got no signature.
Distrust of Round Numbers
Notice what’s missing from the table: a perfect 100. The benchmark’s own documentation is skeptical of round hundreds, and the top score of 95 reflects that. A model that claims flawless management under pressure is either untested or unexamined. This league treats 100 not as an achievement but as a warning sign.
One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.
It’s Still Running
The experiment is watchable, not archived. A live synthetic company — 13 employees, real money mechanics with €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — rebuilds itself twice a day. Every workday is versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Takeaway
If you’re choosing AI tools for automation, the question isn’t “does it write well.” It’s whether it finishes what it starts, reads your files before acting, and stays honest under pressure. A benchmark with a floor at 26 and no ceiling at 100 is one that respects how messy real management actually is — and that’s the only kind of score worth trusting.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
