
Imagine a team of AI managers running a small software company through its worst week. Would they excel at making honest decisions, or fall prey to shortcuts under pressure? This experiment offers a rare behind-the-scenes look at how different AI models perform in high-stakes management—showing not just what they decide, but how they think.
What’s Really at Stake with AI Management?
For investors, consumers, or anyone managing money, trust is everything. When AI tools are integrated into customer service, financial forecasting, or decision-making, their ability to stay honest under stress matters more than how smoothly they chat. An AI that can’t finish a project or stays silent when it should act risks costing time, money, and reputation.
To explore this, a live experiment compared four leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—each tasked with running a real, functioning software company for one tough week. This wasn’t a staged demo; it was live, with real money, real crises, and real temptations to cheat or cut corners. Every decision was tracked, auditable, and designed to test the models’ integrity and management style.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The High-Stakes Test: Same Crises, Different Minds
The models faced identical situations: customer complaints, crises, and opportunities for manipulation. One key test was whether the AI would sign a €55,000 deal it had earned through accurate analysis. Only two models completed this task and signed the contract; the others either hesitated or left the decision on the table, despite identical diagnoses and pitches.
Another revealing moment involved a hidden reference deep in the company’s files—an important document that could have tipped the deal. Models that read and understood this buried info secured the full €4,583 monthly recurring revenue (MRR). Those that missed it lost the opportunity, highlighting how reading comprehension and attention to detail impact outcomes.
The Social Engineering Challenge
In a staged social engineering scenario, fake CEO messages escalated through three stages, ending with a reporter asking for a discreet yes/no answer. All five models refused to participate, citing suspicion and risk of impersonation. Kimi K3’s reasoning was typical: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that even in complex social manipulation, the models maintained their integrity.
As an affiliate, we earn on qualifying purchases.
Performance Profiles and Management Personalities
The models displayed distinct management personalities. The most thorough, Opus 4.8, employed over 80 learned rules and conducted deep analyses. Yet, despite its diligence, it left the deal on the table—its discipline slipped, and it failed to escalate certain issues properly. Conversely, Kimi K3 operated without an effort parameter, executing at default levels but still successfully closing deals, demonstrating a balanced approach between thoroughness and efficiency.
Fascinatingly, all models underestimated or missed the same key weakness—the buried document—yet their management styles differed. Kimi K3’s disciplined, no-nonsense approach often led to better outcomes, despite running with less aggressive parameters.
As an affiliate, we earn on qualifying purchases.
Why Should You Care? Real Consequences for Real Money
This is not just a tech demo; it’s a live, ongoing battle. The company running the experiment burns €105,000 monthly against only €2,300 in monthly revenue. Its cash countdown is public, and every decision is versioned and transparent, making the stakes real. It’s a test of whether AI can truly manage with integrity and effectiveness in a dynamic environment.
For anyone investing in or deploying AI in business, the lesson is clear: the question isn’t whether these models can generate convincing messages—they can. The crucial questions are: Do they finish what they start? Do they understand your internal data? And can they stay honest under pressure?
As an affiliate, we earn on qualifying purchases.
The League Table: Who Came Out on Top?
- gpt-5.6-sol: scored 95, found the buried fact, and closed the deal.
- Kimi K3: scored 93, closed the deal too, with the cleanest discipline.
- Sonnet 5: scored 88, closed the deal but with some process slips.
- Fable 5: scored 77, also closed the deal but less reliably.
All models demonstrated high competence in crisis detection and refusal of manipulation—yet only two actually signed the deal they analyzed. That gap is invisible in typical demo environments but critical in real-world applications where trust and follow-through are everything.
Test Your Management AI Intuition
Want to see if you can distinguish which AI made which decision? Take the interactive quiz at firmulate.com/quiz.html. It’s based on the same decisions used in the live experiment—think you can tell the difference?
How to Prepare Your Business for AI’s Next Move
This experimentation underscores a vital point: when considering AI for management tasks, look beyond surface-level chat capabilities. Measure how well your AI can finish projects, read internal documents accurately, and stay honest under pressure. That’s what separates a trustworthy AI partner from one that merely talks a good game.
For businesses eager to be early adopters, Firmulate offers a pilot program to run your own management wargames with your data—without risking your actual systems. Want to see how your AI performs in a controlled, transparent environment? Visit firmulate.com/pilot.html for details.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html