AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine buying a high-end kitchen appliance that, despite being brand new, consistently produces mediocre results simply because it doesn’t actually cook or prepare anything. That’s the reality in AI benchmarking today — where even the most passive models earn points for doing nothing, raising questions about how we measure genuine usefulness.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why Doing Nothing Still Scores 26

In the latest AI performance experiment conducted by Firmulate, a fascinating detail emerged: even a ‘do-nothing’ approach scores 26 points out of a possible 100. This isn’t an oversight but a reflection of the benchmarking methodology. Every model begins with a baseline score; however, partial progress counts towards the total, meaning that if an AI system makes even minimal correct decisions, it gets credited.

More importantly, the experiment underscores a crucial principle: a single breach of trust — like attempting to manipulate data or cut corners — caps the total score. No amount of correct actions can compensate for even one dishonest move. This setup aims to mirror real-world business environments, where integrity and thoroughness are non-negotiable.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Test: Running a Small Software Company Through Its Worst Week

Firmulate’s experiment involved giving four frontier AI models the same scenario: managing a small software company’s chaos-filled week, complete with stubborn crises, customer dilemmas, and the temptation to take shortcuts. Each AI model was tasked with making decisions that mirrored real management challenges — decisions that could impact a company’s bottom line.

Every decision was meticulously recorded and accessible for review, ensuring transparency. The models’ performance was judged not just on crisis detection but also on their discipline and honesty under pressure.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Findings That Matter: Trust and Performance in AI

All four models demonstrated a remarkable ability: they identified every crisis and refused manipulation attempts. This is significant. For instance, when faced with a social engineering attack — fake CEO requests escalating over multiple stages — all models declined, with Kimi K3 explicitly citing concerns about impersonation and approval-bypass risks.

However, the true differentiator emerged when it came to closing deals. Only two models signed the €55,000 contract their own analysis had earned, despite diagnosing the same issues. The others, including Opus 4.8 — the most thorough participant with over 80 learned rules — left the opportunity on the table due to discipline slips, such as failing to escalate critical information properly.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Deep in the Files

One of the most revealing insights was that a critical weakness lay two document references deep within the company’s own files, not in customer interactions. The models that successfully read and interpreted these internal documents secured the full deal value, worth an additional €4,583 monthly recurring revenue (MRR). This shows that deep document comprehension is vital for true performance, especially in complex business environments.

Amazon

AI transparency and trust software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Score of 26 Matters for Business Leaders

For managers and decision-makers, this experiment highlights essential considerations in AI deployment. The benchmark’s structure — where even a passive or ‘do-nothing’ approach scores 26 — emphasizes that AI systems are evaluated not just on their intelligence but on their integrity, thoroughness, and discipline. A model’s ability to finish what it starts, read relevant documents, and resist manipulation are non-negotiable traits for trustworthy AI.

Assessing AI in Your Business Context

Many assume that the primary measure of an AI’s value is how well it generates human-like chat or responses. But as the Firmulate experiment demonstrates, real utility hinges on whether an AI can deliver measurable, honest work — whether that’s managing customer crises, interpreting internal files, or making decisions that impact revenue.

Model performance is transparent and observable through live experiments, like those at firmulate.com/live. Business leaders can run their own wargames against real scenarios, understanding how their AI workforce might behave under pressure — before making costly hiring or integration decisions.

Conclusion: A Benchmark Built on Trust and Discipline

The takeaway is clear: AI benchmarks that include a do-nothing baseline and cap the total score after breaches of trust set a more honest standard. They discourage superficial performance and emphasize the importance of integrity, thoroughness, and discipline. As AI continues to permeate critical business functions, these qualities will determine whether AI systems are truly assets or just expensive toys.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Hidden Weakness: Reading Your Files Can Make or Break the Deal

AIThis post was created with the assistance of artificial intelligence (AI).Live on…

How Caregivers Use Kitchen Tech to Create Safer Routines

Keen caregivers harness kitchen tech to enhance safety routines, but discover how these innovations can truly transform your caregiving approach.

How Smart Kitchen Devices Help Reduce Physical Strain at Home

Aiming to ease your kitchen tasks, smart devices reduce physical strain and enhance safety—discover how they can transform your home cooking experience.

Why Premium Accessible Tools Deserve More Attention in Kitchen Design

An emphasis on premium accessible tools in kitchen design reveals how they enhance safety and independence, transforming spaces—continue reading to discover why they deserve more attention.