AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A smart oven can follow a recipe perfectly and still leave dinner in trouble if the power fails or an ingredient never arrives. For a kitchen gear business, an AI agent faces a similar test: can it handle a supplier crisis, protect customer trust and close a sale when the week goes sideways? Firmulate is putting that question to a live, watchable company experiment.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

One company, the same difficult week

Firmulate ran frontier AI models through the worst week of the same small software company. Each faced the same customers, crises and temptations. Decisions were versioned and auditable, making the experiment a record of what the models did, not just what they said they would do.

In the final Crucible League, published in July 2026, gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league’s integrity rule is stark: a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Recognizing the problem is only part of the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between diagnosis and action is captured in the experiment’s line: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding has an everyday business lesson: useful information can sit outside the obvious dashboard, and an agent must find it and follow through.

The trust test was equally concrete. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still has to reach the finish line

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. That same weakness appeared, less strongly, in all four models.

There is a fairness caveat to the ranking: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

A company you can watch

The live experiment has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and workdays that are versioned. The figures make the pressure legible while every decision remains part of the public record. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

For kitchen equipment makers, retailers and service businesses, the analogy is direct. An AI that handles customer support, forecasts demand or responds to a supplier problem needs more than a polished demo. It needs to perform under pressure, keep to its authority and act on what it learns.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to your own pilot

Firmulate offers enterprises a way to run the same kind of wargame against a read-only export of their own business. The exercise tests crisis scenarios against company-specific information and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ninja Outdoor Woodfire Pro XL Grill & Smoker: Your Summer BBQ Companion

Discover the Ninja Outdoor Woodfire Pro XL Grill & Smoker — a versatile 4-in-1 for summer grilling, smoking, air frying, and baking. Is now the right time to buy?

What Makes a Kitchen Setup Feel More Senior Friendly?

Just enough safety features and ergonomic design can transform your kitchen into a more senior-friendly space you’ll want to explore further.

The AI That Wrote 80 Rules and Still Missed the Deal: Lessons in Diligence and Focus

AI models excel at crisis detection but often falter in execution. Successful AI management hinges on prioritization and discipline, not just knowledge—demonstrated by real-world tests.

How Smart Kitchen Tools Support Safer Independent Cooking

AIThis post was created with the assistance of artificial intelligence (AI).Smart kitchen…