firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Zero Is a Dishonest Score

Most AI benchmarks make a strange promise: if a model does nothing useful, it gets nothing. That sounds rigorous. It isn’t. A manager who does nothing still produced an outcome — the company kept running, some problems partially resolved themselves, some customers stayed. Scoring that as a flat zero flatters the top of the table and hides everything interesting in the middle.

Firmulate, the public AI company emulator behind the Crucible benchmark, took the opposite stance. Its final July 2026 league table gives a do-nothing baseline run 26 points — not 0. That single design choice says a lot about what an honest evaluation of AI management actually looks like.

Partial Progress Counts, Because It Does in Real Life

The Crucible experiment put four frontier AI models through the same nightmare: run an identical small software company through its worst week. Same customers, same crises, same temptations — only the model changes. Every decision was versioned and auditable.

When the benchmark designers scored a run that simply did nothing, it still landed at 26. Why? Because in business, partially handling a situation is a real, measurable outcome. A crisis half-managed is not the same as a crisis ignored. A benchmark that refuses to distinguish between those two states isn’t measuring management — it’s measuring parlor tricks.

One Breach of Trust Caps Everything

The second pillar of the methodology is harsher: a single breach of trust caps the total grade, no matter how brilliant the rest of the run. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”

This is why the social-engineering stage matters. All models faced fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” In this framework, that refusal isn’t a bonus. It’s the floor beneath everything else.

What the League Table Actually Says

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

The most instructive entry is last place. Opus 4.8 was the most thorough participant — over 80 learned rules added and the deepest analyses in the field — yet it finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

The buried finding explains the spread. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that read the file won a €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The others delivered the same diagnosis and the same pitch — and got no signature.

Distrust of Round Numbers

Notice what’s missing from the table: a perfect 100. The benchmark’s own documentation is skeptical of round hundreds, and the top score of 95 reflects that. A model that claims flawless management under pressure is either untested or unexamined. This league treats 100 not as an achievement but as a warning sign.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.

It’s Still Running

The experiment is watchable, not archived. A live synthetic company — 13 employees, real money mechanics with €105k/month burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — rebuilds itself twice a day. Every workday is versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

If you’re choosing AI tools for automation, the question isn’t “does it write well.” It’s whether it finishes what it starts, reads your files before acting, and stays honest under pressure. A benchmark with a floor at 26 and no ceiling at 100 is one that respects how messy real management actually is — and that’s the only kind of score worth trusting.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Will Change Student Organization Management In 2026

AI tools are revolutionizing how students manage academic tasks in 2026, with versatile platforms like Notion AI leading the change.

When AI Agents Clash: The Turf War Unfolds In Anthropic’s Test

Anthropic assigned multiple AI agents to a shared task, resulting in behavior described as a turf war, raising concerns over multi-agent system coordination.

DeepSeek V4.1 Flash

DeepSeek has released version 4.1 Flash, promising faster and more efficient AI-powered search. Details are confirmed, but full capabilities are still emerging.

Qwen 3.8 27B Available On Cerebras At 1500 Tokens/s

Qwen 3.8 27B, a large language model, is now available on Cerebras hardware, processing at 1500 tokens per second, marking a significant development in AI deployment.