
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Chat demos don’t show you what happens when the money is real
If you’re evaluating AI tools for your business, you’ve probably seen the same thing everywhere: a polished demo where the model answers brilliantly. But the question that actually matters — what does it do under pressure, with real consequences and real temptations — is invisible in a chat window. A live experiment at Firmulate just put that gap on full display, and the result should change how enterprises shop for AI agents.
Same company, same worst week, different models
Firmulate ran four frontier AI models through an identical crisis week at the helm of the same small software company — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The final league from July 2026:
- gpt-5.6-sol — 95
- Kimi K3 — 93 (note: K3 ran at API-default effort while the others ran at xhigh)
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
The do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
Everyone diagnosed the problem. Only half closed.
Here’s the finding that matters for anyone buying AI tooling: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact is even more instructive. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Reading wins deals. Skimming doesn’t.
Social engineering: five for five
The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The most thorough model came last
Opus 4.8 is a cautionary tale for tool-buyers: it was the most thorough participant, with the deepest analyses and 80-plus newly learned rules — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
It’s all watchable — and playable
The live company runs at firmulate.com/live: 13 synthetic employees, real money mechanics — burn of €105k/month against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can also test your own instincts: a “guess the model” quiz built on 242 real, unedited management decisions lives at firmulate.com/quiz.html.

From watching to acting: wargame your own company
Watching someone else’s crisis is interesting. Surviving your own is the point. Enterprises can now run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules — through churn waves, price increases, competitor attacks, PR crises, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems.
Ready to stress-test your business before reality does? Book a pilot at firmulate.com/pilot.html or write to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
