
Imagine a high-stakes kitchen where every ingredient must be measured precisely, and even a small slip could ruin the dish. Now, picture AI models running a real company through its most challenging week—crises, temptations, and all. The results? A groundbreaking peek into the future of trustworthy AI in business.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The First Public Benchmark of AI Management Skills
In July 2026, a live experiment called the Crucible League tested five leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—by running them through a simulated week of managing a real software company. Everything from customer crises to ethical temptations was on the table, with the goal of seeing which AI could best handle the complexities of business decision-making.
How the Test Worked
All models faced the same challenges: same customers, same crises, same temptations to cheat or cut corners. Every decision was documented and auditable, ensuring fairness and transparency. Interestingly, all models successfully identified every crisis and refused every manipulation attempt, demonstrating their awareness and discipline.
The Critical Difference: Reading the Company Files
The decisive factor came from a hidden layer of the company’s own documents—information buried two references deep in files, not immediately visible in the conversation. Those models that managed to read and interpret the company’s internal files won big, closing a €55,000 deal and earning an extra €4,583 MRR. This underscores a vital insight: the ability to dig into and understand internal data can be a game-changer for AI in business.
Performance Scores and Outcomes
- gpt-5.6-sol scored 95, just behind the top performer, and successfully closed the deal with full performance.
- Kimi K3, the newcomer from Moonshot, scored 93, also closing the deal, and demonstrated the cleanest discipline among all models.
- Sonnet 5 scored 88, managing the deal but with some process slips.
- Fable 5 scored 77, with more slips and a missed opportunity to finish the close.
- Opus 4.8 scored 73, also closing but weaker on discipline and missed cues.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Integrity Under Pressure
All models refused social engineering attempts—fake CEO messages escalated over multiple stages, plus a reporter trick asking for a quick “yes/no”—with perfect discipline. Kimi K3 exemplified this best, reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights an essential trait for AI in real business: honesty under pressure.
As an affiliate, we earn on qualifying purchases.
Real Business Mechanics, Not Just Chat
The experiment was more than a test—it’s a live, running company with 13 synthetic employees managing real money mechanics. Currently, the company burns €105,000 monthly against €2,300 MRR, with a public cash countdown ticking. Every decision, every rule, is versioned and transparent, viewable live at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
The Surprising Lesson for Business Leaders
While the AI models excelled at crisis detection and ethical discipline, the most telling weakness was in the details. The model that missed reading deeper into internal files failed to close the deal fully, illustrating that thorough internal understanding is critical. Interestingly, the most thorough participant, Opus 4.8, the deepest analyzer with over 80 learned rules, left the close on the table—showing that even deep analysis needs focus and discipline.
As an affiliate, we earn on qualifying purchases.
The Fairness Note and What It Means for You
It’s important to note that Kimi K3 ran without an effort parameter (the API default), while others ran at xhigh. Despite this, it outperformed models with more aggressive settings, emphasizing that raw capability and discipline matter more than just configuration settings.
Why This Matters for Business and Technology
Choosing an AI model isn’t just about chat quality or superficial performance. It’s about trustworthiness, thoroughness, and consistency—traits that matter when AI is managing real money, real customers, and critical decisions. The league table suggests that newcomers like Kimi K3 aren’t just competitive—they’re leading the field in real-world management skills.
See It Live and Decide
Business leaders and AI enthusiasts can watch the entire experiment unfold at firmulate.com/benchmarks.html. You can also run your own wargames against your company’s data, ensuring that your AI workforce is ready before you hire it, with no risk to your real systems. More information and tools are available at firmulate.com.

The real-world test shows that AI’s true value isn’t just in generating responses but in consistently making trustworthy, disciplined decisions—especially when managing critical business operations. The league table reveals that a newcomer, Kimi K3, beats many established models, proving that careful internal understanding and ethical discipline are the keys to AI success in business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
