
Coding skill is not management skill
For readers choosing AI tools and automation platforms, the familiar leaderboards answer only part of the buying question. They can show whether a model writes strong code or produces a persuasive response. They do not show whether an agent can triage competing emergencies, follow through under capacity pressure, or tell the board an uncomfortable truth.
That measurement gap matters as AI moves from the chat window into the CRM, support queue and financial forecast. A fluent answer may be useful, but a company needs something harder: decisions that remain coherent across days, withstand manipulation and end in completed work.
Firmulate, an AI company emulator, is turning that distinction into a live, watchable experiment. Its proposition is refreshingly direct: measure management quality, not chat quality.
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every model
Firmulate gave frontier models the same small software company and sent each through its worst week. The customers, crises and temptations remained constant; only the model changed. Every decision was versioned and auditable.
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a hard boundary around trust: a single breach capped the total because “no amount of good work outweighs a breach of trust.”
Those results are not merely another model ranking. They expose the distance between recognizing a business problem and resolving it. Every model identified every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
The decisive fact was not in the obvious place
The deal also tested whether an agent would investigate before acting. The decisive weakness in a competitor was buried two document references deep inside the company’s own files, rather than presented in the customer event. Models that found and used that information won the deal at full price, worth +€4,583 MRR.
This is exactly the kind of distinction conventional evaluations often miss. In business, the best response is not necessarily generated from the most visible prompt. It may depend on reading the existing material, connecting a customer request to an internal fact and carrying that advantage all the way through a commercial process.
Pressure tested honesty as well as competence
The experiment subjected the models to fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result should interest any organization considering agents with access to sensitive operational context. The relevant safety question is not simply whether a model knows a policy. It is whether the agent keeps following it while an apparent executive applies pressure or an outsider offers an informal shortcut.
Thoroughness did not guarantee completion
Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, though less strongly.
The lesson is uncomfortable for anyone who equates longer reasoning with better management. Analysis creates value only when it becomes an authorized, completed action. Persistence also needs boundaries: when a route is locked, a capable manager escalates rather than repeatedly pushing the same door.
The comparison includes an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. Readers can examine the full benchmark results with that difference in mind.

A new curriculum for business agents
Scenario names such as churn wave, price increase, downround and PR crisis point toward a more useful curriculum for enterprise AI. These exercises test judgment across connected events: what the agent notices, what it verifies, what it refuses and whether it finishes.
The live company makes those stakes concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays turn automation from a polished demonstration into an observable operating record.
There are also 242 real, unedited management decisions behind Firmulate’s guess-the-model quiz. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.
Before hiring an AI workforce, buyers should ask more than whether a model can produce the right answer. The sharper questions are whether it reads the files, closes the loop, respects authority and remains honest when the week goes badly. That is not chat quality. It is management quality—and it deserves its own category.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html