
When it comes to AI tools transforming your business operations, the focus often lands on how well they generate content or engage with customers. But in high-stakes scenarios—like managing crises or closing deals—these models are tested far beyond their conversational skills. The real question is: can they finish what they start? A groundbreaking live experiment from Firmulate reveals surprising insights about AI decision-making under pressure, showing that visible chat quality isn’t the full story.
The Business-Ready Test: More Than Just Chat
In a recent live experiment, four leading AI models were tasked with running a real software company through its worst week—handling crises, resisting manipulation, and closing a deal worth €55,000. This setup, which mimics real business decision-making, shatters the common assumption that chat demos alone can measure AI readiness. Instead, it exposes something more crucial: the ability to follow through on decisions, read internal files, and maintain discipline under pressure.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Crucible of Crises and Manipulation
Every model successfully identified each crisis and refused attempts at manipulation, including staged social engineering attacks involving fake CEO messages and reporter tricks. The models’ refusal was consistent and reasoned, with Kimi K3 explicitly treating suspicious requests as possible impersonation. This indicates a robust capacity to resist deception—a vital trait for trustworthy AI in business settings.
The Hidden Weakness: Internal File Reading
While all models demonstrated vigilance externally, the decisive difference came from how they utilized internal company documents. The winning models—gpt-5.6-sol and Kimi K3—read deeper into the company’s files, uncovering a buried reference critical to closing the deal. The models that examined this internal information secured the €55,000 deal, equating to an additional €4,583 in Monthly Recurring Revenue (MRR). The loser, Opus 4.8, identified the problem but failed to act on it, leaving the deal on the table.
The Discipline of Follow-Through
The experiment highlights a mismatch between perceived chat quality and actual business performance. For example, Opus 4.8, despite being the most thorough, left the deal unexecuted—its discipline slipping in favor of isolated rule-based checks rather than escalating or executing the final steps. Meanwhile, Kimi K3 ran without an effort parameter, which affected its finisher capabilities but still managed to close the deal, demonstrating that discipline and focus are critical for meaningful work.
Why This Matters for Business Automation
Today’s AI tools are often judged by their conversational prowess—how well they can generate human-like responses. But as this experiment shows, the true measure of an AI’s business utility lies in its ability to follow through on complex tasks, read critical internal information, and resist manipulation under pressure. These are invisible qualities that don’t show up in demos or chat scores but are essential for real-world deployment.
Putting AI to the Test Before Hiring
Firmulate offers a live, watchable environment where companies can run their own ‘wargames’ against AI models—testing decision-making, discipline, and resilience in a risk-free simulation. This approach provides a more truthful measure of an AI’s readiness to replace or augment human teams, avoiding costly missteps from over-reliance on superficial chat capabilities.
The Bottom Line
As AI continues to integrate into core business functions, the question isn’t just about how well it talks but whether it can deliver consistent, trustworthy results. The Firmulate live experiment underscores the importance of testing AI’s ability to execute, read internal context, and resist manipulation—traits that are invisible in demos but vital for success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html