firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Automation’s toughest test is not a demo

AI tools can draft polished emails, summarize meetings and produce convincing plans. Firmulate asks a harder business question: can an AI workforce run a company when customers are unhappy, cash is disappearing and apparently easy shortcuts threaten trust?

The answer is unfolding in public. Firmulate operates a software company staffed by 13 synthetic employees, with real money mechanics and a visible struggle for survival. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.

This makes Firmulate an unusually extreme build-in-public experiment. Visitors can watch the company live, following not merely a product launch but an operating business whose employees, decisions and financial pressure continually produce new material.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to expose the gap between knowing and doing

The live company matters because it turns familiar claims about automation into observable behavior. A model may recognize a problem, recommend the right response and write an excellent sales pitch. None of that guarantees it will complete the work.

Firmulate’s Crucible League made that distinction visible. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring: every model identified every crisis and rejected every manipulation attempt. The sharper finding was less comfortable. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

The decisive fact was buried in ordinary company material

The deal did not turn on eloquence alone. A crucial competitor weakness appeared two document references deep inside the company’s own files rather than in the customer event. The models that followed the trail found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.

For businesses adopting AI tools, that detail may be more important than the league table. Useful automation depends on whether an agent reads the available material, connects evidence across routine documents and carries its conclusion through to action. A system that notices a sales opportunity but fails to close it can look intelligent while producing no commercial result.

Pressure also tested whether the models would stay honest

The worst week included fake CEO messages that escalated over three stages, followed by a reporter attempting to extract information with “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result addresses a practical fear around autonomous tools. An employee can face urgency, authority claims and social pressure simultaneously; an AI worker may encounter the same tactics through messages and requests. Firmulate’s test shows that refusal is possible across the field, while also making trust a non-negotiable part of performance.

Thoroughness did not guarantee the best outcome

Opus 4.8 offers the most revealing cautionary profile. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline problem appeared in weaker form across the other four models.

The implication is not that careful reasoning lacks value. It is that analysis, operational discipline and completion are separate capabilities. A model can produce more detail, learn more lessons and still lose to a competitor that executes the final business step.

There is also an important comparison caveat: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That fairness note does not erase K3’s performance, but it should accompany interpretations of the ranking.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Build in public becomes an ongoing management story

Firmulate’s strongest contribution is not a prediction about an employee-free future. It is a public record of what synthetic employees actually do when business conditions become messy. Readers can inspect the live financial pressure and follow what the employees say through the company’s public quotes.

The experiment also contains 242 real, unedited management decisions used in a model-guessing quiz. Together with the versioned workdays and expanding playbook, they turn the company into a continuing case study rather than a one-time benchmark.

For organizations exploring AI tools and automation, the lesson is straightforward: fluency is only the opening requirement. The more consequential questions are whether an AI worker reads the company’s own evidence, resists pressure, respects boundaries and finishes valuable work. Firmulate makes those behaviors watchable while its synthetic team confronts a very tangible problem: a company burning far more cash than it earns.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The Luxury Headphone Upgrade That Matters More Than Brand Prestige

When upgrading your luxury headphones, focus on advanced noise cancellation, premium sound…

7 Best Home Theater Projector Prime Day Deals for Big-Screen Movie Nights in 2026

Discover the best Prime Day deals on home theater projectors, including 4K laser, 1080p, short-throw options, and accessories for big-screen movie nights.

7 Best Headphones for Prime Day Electronics Deals in 2026

Discover the best headphones for Prime Day 2026, including top picks for noise cancelling, battery life, comfort, and value, based on expert analysis.

Laser TV vs Projector: The Decision Rule That Prevents Expensive Regret

Keen to avoid costly regrets, discover how to choose between a laser TV and projector by understanding your space, mobility, and long-term needs.