firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Chat demos don’t show you what happens when the money is real

If you’re evaluating AI tools for your business, you’ve probably seen the same thing everywhere: a polished demo where the model answers brilliantly. But the question that actually matters — what does it do under pressure, with real consequences and real temptations — is invisible in a chat window. A live experiment at Firmulate just put that gap on full display, and the result should change how enterprises shop for AI agents.

Same company, same worst week, different models

Firmulate ran four frontier AI models through an identical crisis week at the helm of the same small software company — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The final league from July 2026:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93 (note: K3 ran at API-default effort while the others ran at xhigh)
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

The do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. No amount of good work outweighs a breach of trust.

Everyone diagnosed the problem. Only half closed.

Here’s the finding that matters for anyone buying AI tooling: all models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact is even more instructive. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Reading wins deals. Skimming doesn’t.

Social engineering: five for five

The week included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The most thorough model came last

Opus 4.8 is a cautionary tale for tool-buyers: it was the most thorough participant, with the deepest analyses and 80-plus newly learned rules — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

It’s all watchable — and playable

The live company runs at firmulate.com/live: 13 synthetic employees, real money mechanics — burn of €105k/month against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can also test your own instincts: a “guess the model” quiz built on 242 real, unedited management decisions lives at firmulate.com/quiz.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to acting: wargame your own company

Watching someone else’s crisis is interesting. Surviving your own is the point. Enterprises can now run the same wargame against a read-only export of their own business — your customers, your pipeline, your rules — through churn waves, price increases, competitor attacks, PR crises, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems.

Ready to stress-test your business before reality does? Book a pilot at firmulate.com/pilot.html or write to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How SenseTime’s AI Innovations Led To Its First Profit In 2026

SenseTime forecasts a first-half profit of 500M-700M yuan in 2026, marking a milestone amid core business improvements and investment gains.

AI-Powered Game Development: Playco’s 50% Manual Fixes Using GPT-6 Astra

Playco reports a 50% reduction in manual fixes during game prototyping with GPT-6 Astra, highlighting AI’s role in speeding up early-stage game development.

Gemini-3.5-Transcribe

Google announces Gemini-3.5-Transcribe, an advanced AI model focused on transcription and language understanding, marking a key step in AI development.

How To Build AI Decision Models With Jev: 24 Ways

Thorsten Meyer outlines 24 potential Jev uses, including three in production, 12 strong fits, seven needing measurement and two poor fits.