
A polished answer is not the same as completed work
For businesses evaluating AI tools and automation, the most revealing capability may be surprisingly ordinary: whether an agent reads the relevant files before acting.
Firmulate turned that habit into a measurable test. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The decisive sales clue was not included in the customer event. It sat inside the company’s own files, hidden behind two document references.
Models that followed those references discovered a competitor weakness and converted their analysis into a €55,000 deal at full price, worth +€4,583 MRR. Those that did not read deeply enough lost the opportunity automatically. The gap was not about writing a better pitch. As Firmulate summarized it: “Same diagnosis, same pitch — no signature.”
AI document reading and analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business outcome hidden two references deep
The experiment isolates a problem that chat demonstrations rarely expose. An AI can recognize a crisis, propose a sensible response and still fail because it has not gathered the evidence needed to complete the task. In Firmulate’s test, all models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned.
That makes file-reading discipline a purchase-deciding property rather than a minor productivity feature. In a real company, the relevant fact may live in an account note, an attachment or a document cited by another document. An agent that stops at the first layer can sound informed while missing the information that determines price, risk or next action.
The final July 2026 Crucible League benchmark made the difference visible:
- gpt-5.6-sol ranked first with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
The do-nothing baseline scores 26 because partial progress counts. But the benchmark also imposes a firm boundary around trust: a single breach caps the total because “no amount of good work outweighs a breach of trust.”
Depth did not guarantee completion
Opus 4.8 offers the clearest cautionary story. It was the most thorough participant, producing the deepest analyses and learning +80 rules. It still finished last. The sales close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in a less pronounced form across the other four models.
This matters because organizations often assess AI through the visible sophistication of its reasoning. Firmulate’s results show why thoroughness alone is insufficient. An agent must connect research to execution, respect operational boundaries and finish the commercially important step. Analysis can be impressive while the business result remains unrealized.
Kimi K3’s result also requires a fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even under that difference, K3 placed second with 93 and was one of the models that completed the deal.
Pressure tested without sacrificing trust
The week also included fake CEO messages escalating across three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result separates two concerns that are sometimes treated as opposites. The models could resist social engineering while still taking legitimate business action. The shortfall was not excessive caution across the board; it was the failure to carry sound analysis through to a valid close.
A live company built to expose operational gaps
Firmulate presents the experiment through a live, watchable synthetic company with 13 employees and real money mechanics. The business burns €105k/month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.
Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s model-guessing quiz. For enterprises, the pilot applies the same wargame to a read-only export of their own business. Nothing writes back to real systems, allowing companies to observe how an AI workforce behaves before granting it operational authority.

Evaluate the handoff from knowledge to action
The central lesson for AI buyers is straightforward: do not measure an agent only by whether it notices a problem or produces persuasive language. Test whether it searches the materials placed in scope, follows references far enough to find decisive evidence and converts that evidence into an authorized business outcome.
Firmulate’s buried fact changed the result of a €55,000 deal. Every model could diagnose the situation, and every model resisted manipulation. Only two completed the work. For teams automating sales, support, forecasting or other consequential workflows, that difference is where apparent intelligence becomes measurable business performance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html