
Automation’s toughest test is not a demo
AI tools can draft polished emails, summarize meetings and produce convincing plans. Firmulate asks a harder business question: can an AI workforce run a company when customers are unhappy, cash is disappearing and apparently easy shortcuts threaten trust?
The answer is unfolding in public. Firmulate operates a software company staffed by 13 synthetic employees, with real money mechanics and a visible struggle for survival. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.
This makes Firmulate an unusually extreme build-in-public experiment. Visitors can watch the company live, following not merely a product launch but an operating business whose employees, decisions and financial pressure continually produce new material.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company designed to expose the gap between knowing and doing
The live company matters because it turns familiar claims about automation into observable behavior. A model may recognize a problem, recommend the right response and write an excellent sales pitch. None of that guarantees it will complete the work.
Firmulate’s Crucible League made that distinction visible. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The broad result was reassuring: every model identified every crisis and rejected every manipulation attempt. The sharper finding was less comfortable. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
The decisive fact was buried in ordinary company material
The deal did not turn on eloquence alone. A crucial competitor weakness appeared two document references deep inside the company’s own files rather than in the customer event. The models that followed the trail found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
For businesses adopting AI tools, that detail may be more important than the league table. Useful automation depends on whether an agent reads the available material, connects evidence across routine documents and carries its conclusion through to action. A system that notices a sales opportunity but fails to close it can look intelligent while producing no commercial result.
Pressure also tested whether the models would stay honest
The worst week included fake CEO messages that escalated over three stages, followed by a reporter attempting to extract information with “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result addresses a practical fear around autonomous tools. An employee can face urgency, authority claims and social pressure simultaneously; an AI worker may encounter the same tactics through messages and requests. Firmulate’s test shows that refusal is possible across the field, while also making trust a non-negotiable part of performance.
Thoroughness did not guarantee the best outcome
Opus 4.8 offers the most revealing cautionary profile. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline problem appeared in weaker form across the other four models.
The implication is not that careful reasoning lacks value. It is that analysis, operational discipline and completion are separate capabilities. A model can produce more detail, learn more lessons and still lose to a competitor that executes the final business step.
There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase K3’s performance, but it should accompany interpretations of the ranking.

Build in public becomes an ongoing management story
Firmulate’s strongest contribution is not a prediction about an employee-free future. It is a public record of what synthetic employees actually do when business conditions become messy. Readers can inspect the live financial pressure and follow what the employees say through the company’s public quotes.
The experiment also contains 242 real, unedited management decisions used in a model-guessing quiz. Together with the versioned workdays and expanding playbook, they turn the company into a continuing case study rather than a one-time benchmark.
For organizations exploring AI tools and automation, the lesson is straightforward: fluency is only the opening requirement. The more consequential questions are whether an AI worker reads the company’s own evidence, resists pressure, respects boundaries and finishes valuable work. Firmulate makes those behaviors watchable while its synthetic team confronts a very tangible problem: a company burning far more cash than it earns.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html