AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine deploying new AI tools in your finances or investments, hoping they’ll outperform established players. The question isn’t just about language or chatter — it’s whether these AI systems can consistently deliver results, resist manipulation, and stay honest when stakes are high. Recent experiments in AI-driven management reveal that even newcomers can beat veterans, with profound implications for how we trust and adopt automation in our financial lives.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

AI Models Face the Reality Test: A Live Business War Game

In a groundbreaking live experiment, four leading AI models were put through the same brutal week as a real small software company. Every decision, crisis, and temptation was identical across all models, creating a level playing field for evaluation. The goal was simple yet critical: can these AI systems not only diagnose and analyze but also act decisively — and honestly — in a high-pressure environment?

All four models demonstrated impressive awareness by identifying every crisis and resisting attempts at manipulation. Whether facing fake CEO messages or external reporters pushing for approvals, each AI refused to be duped. This ensures that AI-driven management tools won’t blindly comply with unethical requests — a vital trait for safeguarding business integrity.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Winner: Moonshot’s Kimi K3

While all four models showed strong crisis detection and honesty, one stood out by completing the entire management task successfully. Kimi K3, developed by Moonshot, scored a 93 out of 100 in the final Crucible league, only slightly behind the top-scoring GPT-5.6-sol at 95. Its performance wasn’t just about avoiding pitfalls but also about seizing opportunities—most notably, uncovering a hidden document that led to closing a lucrative €55,000 deal, translating into an extra €4,583 in monthly recurring revenue.

Remarkably, this newcomer beat three of the four Western frontier models, including some with longer track records, in a fair, live setting. The league leaderboard underscores an important shift: the top AI models are closing the gap faster than many anticipated. For investors and business leaders, this means that choosing an AI partner today involves more than just chat scores. It’s about real-world performance, discipline, and integrity under pressure.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Did the Models Need to Win? Reading Deeper Than Surface

The key to K3’s success lay in its ability to dig deeper into the company’s internal files—two document references beneath the surface—allowing it to make informed, profitable decisions. Models that merely skimmed or relied on superficial cues missed this critical insight. As a result, K3’s analysis turned into a full-price deal, securing additional revenue and demonstrating the importance of thorough document comprehension in AI management systems.

Amazon

AI cybersecurity for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Manipulation

Another significant finding was all models’ refusal to be manipulated through fake CEO messages and staged reporter requests. The experiment staged escalating tactics, yet none succeeded — a testament to the models’ built-in safeguards against deception. Kimi K3 explained its stance clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial for real-world applications where malicious actors may try to exploit AI systems.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Reality of Running AI as a Business

The live company used in this test comprises 13 synthetic employees operating in a real money environment—burning €105,000 each month against a revenue of only €2,300. The setup is real, dynamic, and openly observable at firmulate.com/live. The AI models manage every aspect of the company’s operations, from decision-making to crisis handling, with over 680 self-learned rules continually refined. This transparency provides a rare window into how AI can truly perform in high-stakes settings.

The Lessons for Investors and Managers

  • Performance Matters More Than Promises: The league table shows the real winners and losers, with Kimi K3 just behind the top dog, GPT-5.6-sol. The difference in scores reflects actual decision-making quality, not just language fluency.
  • Deep Reading and Discipline Are Crucial: K3’s success hinged on reading beneath the surface and resisting shortcuts or slips in discipline. The most thorough participant, Opus 4.8, left a deal on the table due to a lapse—highlighting that thoroughness and adherence to protocols matter.
  • Trust Is Built on Consistency and Security: All models refused manipulation attempts, reinforcing that AI systems can be reliable under pressure if properly designed. For finance, this means trusting AI to stay honest and focused on actual work, not just sounding convincing.
  • The Future Is Open and Competitive: The league’s results show that the AI field is evolving fast, and the best choice depends on testing models in your context—no longer a game for chat demos but a real decision-making force.

Fairness and Testing Conditions

It’s important to note that K3 ran without an effort parameter (the API default), while others ran at xhigh. This difference underscores that, even under different conditions, K3’s performance was remarkable.

What’s Next?

For enterprises and investors, the takeaway is clear: testing AI models in realistic, high-pressure environments can reveal their true capabilities. Firms offering such live wargames—like Firmulate—are pioneering new standards for evaluation, moving beyond superficial chat demos to real-world readiness. As AI continues to mature, performance consistency, integrity, and the ability to uncover hidden insights will become the key criteria for trust and success.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Hyperscaler Prisoner’s Dilemma: Why I Keep Buying Nvidia

Analysis of why hyperscaler companies continue purchasing Nvidia stock despite market uncertainties, highlighting the strategic and economic factors involved.

Aktualisierte Sanktionsmeldung: Sudan

FINMA releases an updated sanctions report on Sudan, detailing new restrictions and measures amid ongoing conflicts and regional tensions.

After Macron’s Example, Trump Teams up With PM Modi to Advance Artificial Intelligence

Leveraging Macron’s vision, Trump and Modi unite to revolutionize AI—could this alliance alter the dynamics of global technology and diplomacy?

Crypto and Trump: What His Alleged Bitcoin Price Target Means

Potential shifts in the crypto landscape loom as Trump’s alleged Bitcoin price target emerges, but what could this mean for the future of digital assets?