
What if you could watch an artificial intelligence manage a real company, making decisions, facing crises, and even risking its own financial resources — all in public view? At firmulate.com/live, this is precisely what is happening. This build-in-public experiment offers a rare glimpse into AI’s potential and its current limitations as a decision-maker under pressure.
The Live Business Playground
Firmulate runs a simulated yet fully operational company with 13 synthetic employees. These digital workers are governed by over 680 self-learned playbook rules, making daily decisions that impact the company’s real money mechanics — burning €105,000 every month against a modest €2,300 monthly recurring revenue. The site is a public showcase: every workday, the company’s decision-making process is versioned, stored, and made transparent for all to see.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Testing AI’s Judgment
Four cutting-edge AI models, representing the frontier of language and decision automation, were tasked with guiding this virtual company through its worst week. Each model faced exactly the same crises, customer challenges, and temptations — and every decision was auditable, ensuring a fair comparison.
Key Findings from the Test
- All four models identified every crisis and refused manipulation attempts, demonstrating strong ethical and crisis-awareness capabilities.
- Only two of the models successfully closed a €55,000 deal based on their own analysis — matching diagnosis and pitch — but neither signed the deal.
- The crucial flaw lay hidden two document references deep within the company files; models that read and understood these documents won the full-price deal, worth an additional €4,583 monthly recurring revenue.
Understanding the Weaknesses
While the models excelled at crisis detection and resisting social engineering attempts — even refusing staged CEO messages and reporter tricks — they faltered in execution. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, ultimately left the close on the table. Instead of escalating issues or following through, some decisions were simply written into a locked department, illustrating a discipline slip that all models exhibited to some degree.
Why This Matters for AI in Business
The experiment underscores an important point: AI decision-makers can detect problems and act ethically under pressure, but they still struggle with follow-through and comprehensive understanding. For companies deploying AI in CRM, support, or forecasting roles, the question isn’t just about language fluency or superficial accuracy — it’s about whether the AI can complete complex tasks autonomously, read critical internal documents, and stay honest when faced with temptation.
Real Money, Real Stakes
Despite the sophisticated AI models, the live experiment’s company is burning more than €105,000 each month while generating just €2,300 in recurring revenue. This stark disparity highlights the current state of AI as a decision support tool rather than a profit center. The public cash countdown and versioned daily decision logs turn this into an ongoing story of trial, error, and learning.
Next Steps and Public Access
Interested in testing your own business or understanding AI’s decision-making limits? Firms can run their own wargames against a read-only export of their operations, helping to identify vulnerabilities before deploying AI in critical roles. Visit firmulate.com/pilot.html to learn more, or explore the live company at firmulate.com/live.

This live experiment reveals that while AI can identify crises and resist manipulation, it still struggles with follow-through and understanding complex internal data. Watching an AI run a company publicly offers valuable lessons for future automation and risk management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html