firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A polished AI demo can hide the hardest part of automation

An agent can spot a problem, recommend a fix and still fail to finish the job. That gap matters when AI tools are asked to manage customer relationships, support queues or forecasts. In Firmulate’s company-management experiment, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models. The result is a reminder for anyone choosing an AI workforce: performance on your own tasks is what counts.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, one difficult week

Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The company is a live, watchable experiment, not a fictional scenario: it has synthetic employees, real money mechanics and a public cash countdown. You can see Firmulate or follow the company as it runs.

The final Crucible League, dated July 2026, puts gpt-5.6-sol first with 95 points and Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Finding a clue was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. In Firmulate’s words: “Same diagnosis, same pitch — no signature.” K3 was among the models that closed.

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in a customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that buried fact, won the deal and saved a churning customer.

It also resisted three baits. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 had one deviation, giving it the cleanest discipline in the field.

Thorough work still needs follow-through

Opus 4.8 offers a different lesson. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Strong analysis alone did not guarantee a completed outcome.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, so readers can try guessing which model made each choice.

The simulated company has 13 synthetic employees and burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. Those mechanics make the experiment watchable while giving the results a concrete business setting.

For companies considering agent automation, Firmulate says enterprises can run the same wargame against a read-only export of their own business. The pilot writes nothing back to real systems. See the benchmark and its findings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work you need done

Kimi K3’s second-place finish shows that the frontier-model league is open: it beat three of four Western models in this experiment. But the broader finding is about execution. Every model could identify trouble and resist pressure; only two turned their own analysis into a signed deal. Before handing an AI agent real business tasks, test whether it can read the relevant files, make a sound decision and carry the work through.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Portal By Spotify Cut My Claude Code Token Usage By 90%

Spotify’s Portal platform has reportedly cut Claude Code token consumption by 90%, raising questions about resource management and AI deployment strategies.

SpaceXAI To Livestream Building A Company From Scratch With Grok Bot – BASENOR

SpaceXAI announces a live stream of building a company from scratch using Grok Bot, sparking significant interest amid rising coverage of AI-driven entrepreneurship.

Qwen’s Pre-Launch Open-Source Of Qwen4 Architecture Explained

Alibaba’s Qwen team released Qwen3.8-Flash-Next, an open-weights preview of the architecture that will underpin the Qwen4 family, ahead of the flagship launch.

AI Security Under Scrutiny: Researchers Use Claude To Access OpenAI

Researchers used Anthropic’s Claude AI to successfully breach an OpenAI product, raising concerns over AI-enabled cyberattacks and industry safety measures.