
For anyone who follows a company’s cash, customers or investment case, a polished pitch is not the same as a signed deal. Firmulate’s business experiment puts that distinction under pressure: AI models ran the same small software company through a week of crises, and most identified the opportunity. Only two closed it.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate is a live, watchable experiment in running a synthetic company with AI. Its 13 synthetic employees face real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Workdays are versioned, and the company has accumulated more than 680 self-learned playbook rules.
For its final Crucible League in July 2026, each frontier model ran the same small software company through its worst week. Customers, crises and temptations were held constant; every decision was versioned and auditable. The league ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 fifth at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Recognizing the crisis was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summed up the gap: “Same diagnosis, same pitch — no signature.”
The detail behind the deal matters. The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding suggests that the quality of an AI workforce depends on how it handles a real company’s scattered information and follows through on its own conclusions.
Integrity faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work, unfinished business
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz based on 242 real, unedited management decisions, inviting readers to guess which model made each call.
From watching to a company-specific test
A live synthetic company can show how models behave under shared conditions. An enterprise pilot takes the idea to a company’s own business: Firmulate says it can run the wargame against a read-only export, stage crisis scenarios against that business, and produce a board report with model rankings and weak points in existing playbooks. Nothing writes back to real systems.
That boundary is central for leaders weighing AI in customer management, support or forecasting. A model can spot a problem and still fail to complete the job. A controlled pilot offers a way to examine those decisions against company-specific information before entrusting agents with live work.

Take the next step
To explore a company-specific AI wargame, visit Firmulate’s pilot page or contact contact@firmulate.com. You can also follow the live experiment at Firmulate.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
