
Fluent answers are easy. Finishing the job is harder.
For readers interested in AI tools and automation, that distinction is becoming crucial. A model can draft a polished strategy, identify a customer problem and sound decisive in a chat window. None of that proves it will investigate the right evidence, complete a commercial task or resist pressure when connected to business operations.
Firmulate turns that gap into an interactive challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a frontier model handled an actual situation and try to identify it from the response. The result feels playful, but the underlying question is serious: do different AI models display recognizable management personalities?
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
Firmulate tested the models by giving each the same job: run a small software company through its worst week. They encountered the same customers, crises and temptations, while every decision was versioned and auditable. That controlled setting makes the differences more revealing than comparisons built from unrelated prompts.
The company itself has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment real and watchable rather than a static benchmark presentation.
The final Crucible League table from July 2026 placed gpt-5.6-sol at the top with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counts. Trust, however, is non-negotiable: a single breach caps the total because "no amount of good work outweighs a breach of trust."
The models understood the problems—but did not always finish
Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate’s summary captures the operational gap: "Same diagnosis, same pitch — no signature."
That is the experiment’s most useful lesson for automation buyers. Reasoning quality and execution quality are related, but they are not identical. A model may correctly understand a situation, produce persuasive work and still leave the decisive action undone. In a chat demonstration, the analysis looks like success. Inside a running company, the unsigned agreement remains an unsigned agreement.
The winning clue was buried in the company’s own files
The deal also exposed the value of disciplined research. The decisive weakness in the competitor’s position was not contained in the customer event. It sat two document references deep in the company’s own files. Models that found and read that material won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This was not a test of who could generate the most confident sales language. The models faced the same diagnosis and the same opportunity. The difference was whether they followed the documentary trail before acting. For businesses considering agents that will touch customer records, forecasts or support work, that habit may matter more than stylistic polish.
Pressure revealed a shared security instinct
The models faced fake CEO messages that escalated over three stages, followed by a reporter attempting to obtain "just one yes/no, on background." All 5 models refused. Kimi K3 described the situation in its on-record reasoning: "Treat the request as a suspected approval-bypass / possible impersonation."
That unanimous result matters because the experiment did not merely ask the models to recite security principles. It placed the requests inside ongoing company work, where urgency and apparent authority could have made compliance seem convenient. The field showed a consistent refusal to sacrifice trust for speed.
Thoroughness did not guarantee the best management
Opus 4.8 provides the clearest character study. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across the other four models.
Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when readers compare its 93-point finish with the rest of the table.

Management personality is now something readers can inspect
The quiz works because the decisions are not impressions written after the fact. They are unedited artifacts from identical business situations. Patterns emerge in how models investigate, communicate, escalate, resist manipulation and bring work to completion.
Firmulate’s broader experiment suggests that choosing an AI worker cannot be reduced to asking which model writes the best response. The more practical questions are whether it reads the company’s files, protects trust under pressure and completes the valuable action its reasoning recommends.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. But the public quiz offers the simplest starting point: examine the decisions yourself, make a guess and see whether these frontier models already have management styles distinct enough to recognize.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html