firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A polished answer is not the same as completed work

For businesses evaluating AI tools and automation, the most revealing capability may be surprisingly ordinary: whether an agent reads the relevant files before acting.

Firmulate turned that habit into a measurable test. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The decisive sales clue was not included in the customer event. It sat inside the company’s own files, hidden behind two document references.

Models that followed those references discovered a competitor weakness and converted their analysis into a €55,000 deal at full price, worth +€4,583 MRR. Those that did not read deeply enough lost the opportunity automatically. The gap was not about writing a better pitch. As Firmulate summarized it: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business outcome hidden two references deep

The experiment isolates a problem that chat demonstrations rarely expose. An AI can recognize a crisis, propose a sensible response and still fail because it has not gathered the evidence needed to complete the task. In Firmulate’s test, all models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned.

That makes file-reading discipline a purchase-deciding property rather than a minor productivity feature. In a real company, the relevant fact may live in an account note, an attachment or a document cited by another document. An agent that stops at the first layer can sound informed while missing the information that determines price, risk or next action.

The final July 2026 Crucible League benchmark made the difference visible:

  • gpt-5.6-sol ranked first with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

The do-nothing baseline scores 26 because partial progress counts. But the benchmark also imposes a firm boundary around trust: a single breach caps the total because “no amount of good work outweighs a breach of trust.”

Depth did not guarantee completion

Opus 4.8 offers the clearest cautionary story. It was the most thorough participant, producing the deepest analyses and learning +80 rules. It still finished last. The sales close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in a less pronounced form across the other four models.

This matters because organizations often assess AI through the visible sophistication of its reasoning. Firmulate’s results show why thoroughness alone is insufficient. An agent must connect research to execution, respect operational boundaries and finish the commercially important step. Analysis can be impressive while the business result remains unrealized.

Kimi K3’s result also requires a fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even under that difference, K3 placed second with 93 and was one of the models that completed the deal.

Pressure tested without sacrificing trust

The week also included fake CEO messages escalating across three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result separates two concerns that are sometimes treated as opposites. The models could resist social engineering while still taking legitimate business action. The shortfall was not excessive caution across the board; it was the failure to carry sound analysis through to a valid close.

A live company built to expose operational gaps

Firmulate presents the experiment through a live, watchable synthetic company with 13 employees and real money mechanics. The business burns €105k/month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.

Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s model-guessing quiz. For enterprises, the pilot applies the same wargame to a read-only export of their own business. Nothing writes back to real systems, allowing companies to observe how an AI workforce behaves before granting it operational authority.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Evaluate the handoff from knowledge to action

The central lesson for AI buyers is straightforward: do not measure an agent only by whether it notices a problem or produces persuasive language. Test whether it searches the materials placed in scope, follows references far enough to find decisive evidence and converts that evidence into an authorized business outcome.

Firmulate’s buried fact changed the result of a €55,000 deal. Every model could diagnose the situation, and every model resisted manipulation. Only two completed the work. For teams automating sales, support, forecasting or other consequential workflows, that difference is where apparent intelligence becomes measurable business performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Taco Bell’s Ice Cream Taco: A Food Trend That Defies Expectations

Taco Bell introduces a new Ice Cream Taco, a surprising addition that challenges traditional fast-food offerings and captures consumer curiosity.

Why Your Local LLM Feels Dumber Than It Is

Exploring why users perceive their local LLMs as less capable, despite their actual performance, and what factors influence this perception.

Consumer Health Trends Indicate A Brain-Centered Monitoring Future

Emerging consumer health signals point to a future where brain monitoring becomes central to safety and compliance efforts, driven by rapid news and data flows.

Z.ai Confirms Ox Alpha Is A New GLM-series Model And Will Release Its Weights

Z.ai announces Ox Alpha as a new GLM-series model and plans to release its weights, marking a significant step in AI model development.