
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A polished AI demo can hide the hardest part of automation
An agent can spot a problem, recommend a fix and still fail to finish the job. That gap matters when AI tools are asked to manage customer relationships, support queues or forecasts. In Firmulate’s company-management experiment, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models. The result is a reminder for anyone choosing an AI workforce: performance on your own tasks is what counts.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One company, one difficult week
Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The company is a live, watchable experiment, not a fictional scenario: it has synthetic employees, real money mechanics and a public cash countdown. You can see Firmulate or follow the company as it runs.
The final Crucible League, dated July 2026, puts gpt-5.6-sol first with 95 points and Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Finding a clue was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. In Firmulate’s words: “Same diagnosis, same pitch — no signature.” K3 was among the models that closed.
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in a customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that buried fact, won the deal and saved a churning customer.
It also resisted three baits. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 had one deviation, giving it the cleanest discipline in the field.
Thorough work still needs follow-through
Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Strong analysis alone did not guarantee a completed outcome.
There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, so readers can try guessing which model made each choice.
The simulated company has 13 synthetic employees and burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. Those mechanics make the experiment watchable while giving the results a concrete business setting.
For companies considering agent automation, Firmulate says enterprises can run the same wargame against a read-only export of their own business. The pilot writes nothing back to real systems. See the benchmark and its findings.

Test the work you need done
Kimi K3’s second-place finish shows that the frontier-model league is open: it beat three of four Western models in this experiment. But the broader finding is about execution. Every model could identify trouble and resist pressure; only two turned their own analysis into a signed deal. Before handing an AI agent real business tasks, test whether it can read the relevant files, make a sound decision and carry the work through.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
