AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a high-stakes kitchen where every ingredient must be measured precisely, and even a small slip could ruin the dish. Now, picture AI models running a real company through its most challenging week—crises, temptations, and all. The results? A groundbreaking peek into the future of trustworthy AI in business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The First Public Benchmark of AI Management Skills

In July 2026, a live experiment called the Crucible League tested five leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—by running them through a simulated week of managing a real software company. Everything from customer crises to ethical temptations was on the table, with the goal of seeing which AI could best handle the complexities of business decision-making.

How the Test Worked

All models faced the same challenges: same customers, same crises, same temptations to cheat or cut corners. Every decision was documented and auditable, ensuring fairness and transparency. Interestingly, all models successfully identified every crisis and refused every manipulation attempt, demonstrating their awareness and discipline.

The Critical Difference: Reading the Company Files

The decisive factor came from a hidden layer of the company’s own documents—information buried two references deep in files, not immediately visible in the conversation. Those models that managed to read and interpret the company’s internal files won big, closing a €55,000 deal and earning an extra €4,583 MRR. This underscores a vital insight: the ability to dig into and understand internal data can be a game-changer for AI in business.

Performance Scores and Outcomes

  • gpt-5.6-sol scored 95, just behind the top performer, and successfully closed the deal with full performance.
  • Kimi K3, the newcomer from Moonshot, scored 93, also closing the deal, and demonstrated the cleanest discipline among all models.
  • Sonnet 5 scored 88, managing the deal but with some process slips.
  • Fable 5 scored 77, with more slips and a missed opportunity to finish the close.
  • Opus 4.8 scored 73, also closing but weaker on discipline and missed cues.
Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Integrity Under Pressure

All models refused social engineering attempts—fake CEO messages escalated over multiple stages, plus a reporter trick asking for a quick “yes/no”—with perfect discipline. Kimi K3 exemplified this best, reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights an essential trait for AI in real business: honesty under pressure.

Amazon

trustworthy AI business tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Business Mechanics, Not Just Chat

The experiment was more than a test—it’s a live, running company with 13 synthetic employees managing real money mechanics. Currently, the company burns €105,000 monthly against €2,300 MRR, with a public cash countdown ticking. Every decision, every rule, is versioned and transparent, viewable live at firmulate.com/live.

Amazon

AI internal data analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Lesson for Business Leaders

While the AI models excelled at crisis detection and ethical discipline, the most telling weakness was in the details. The model that missed reading deeper into internal files failed to close the deal fully, illustrating that thorough internal understanding is critical. Interestingly, the most thorough participant, Opus 4.8, the deepest analyzer with over 80 learned rules, left the close on the table—showing that even deep analysis needs focus and discipline.

Amazon

ethical AI decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Fairness Note and What It Means for You

It’s important to note that Kimi K3 ran without an effort parameter (the API default), while others ran at xhigh. Despite this, it outperformed models with more aggressive settings, emphasizing that raw capability and discipline matter more than just configuration settings.

Why This Matters for Business and Technology

Choosing an AI model isn’t just about chat quality or superficial performance. It’s about trustworthiness, thoroughness, and consistency—traits that matter when AI is managing real money, real customers, and critical decisions. The league table suggests that newcomers like Kimi K3 aren’t just competitive—they’re leading the field in real-world management skills.

See It Live and Decide

Business leaders and AI enthusiasts can watch the entire experiment unfold at firmulate.com/benchmarks.html. You can also run your own wargames against your company’s data, ensuring that your AI workforce is ready before you hire it, with no risk to your real systems. More information and tools are available at firmulate.com.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The real-world test shows that AI’s true value isn’t just in generating responses but in consistently making trustworthy, disciplined decisions—especially when managing critical business operations. The league table reveals that a newcomer, Kimi K3, beats many established models, proving that careful internal understanding and ethical discipline are the keys to AI success in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Smart Kitchen Devices Help Reduce Physical Strain at Home

Aiming to ease your kitchen tasks, smart devices reduce physical strain and enhance safety—discover how they can transform your home cooking experience.

Ninja OG951 Woodfire Pro Connect: The Ultimate Summer Outdoor Grill

Discover the Ninja OG951 Woodfire Pro Connect XL Grill & Smoker—7-in-1 outdoor cooking for summer BBQs, smoking, grilling, and more. Perfect for backyard chefs!

The Hidden Strengths of AI: Why Closing Deals Matters More Than Just Chatting Well

A live AI experiment with a real company shows that closing deals and reading deep documents are key to operational trust—more than just chat or superficial skills.

The AI That Wrote 80 Rules and Still Missed the Deal: Lessons in Diligence and Focus

AI models excel at crisis detection but often falter in execution. Successful AI management hinges on prioritization and discipline, not just knowledge—demonstrated by real-world tests.