AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine trusting an AI to handle your most sensitive business decisions—only to find it refuses to sign a simple deal or read critical files. For couples and companies alike, trust is everything. But how do we measure it?

Before you orderOffer from Amazon

Get gifts for the two of you delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Tests Trust and Discipline

Recently, a groundbreaking experiment called the Firmulate AI benchmark put four advanced AI models through a simulated week of running a small software company. This wasn’t about chatty responses or clever tricks; it was about managing crises, reading critical documents, and making honest decisions under pressure.

The Do-Nothing Baseline: Why It Scores 26 Points

One striking finding emerged: even a ‘do-nothing’ AI, which merely observes and refrains from acting, scores 26 out of 100 points. This score exists because partial progress—like reading a file or recognizing a crisis—counts towards the total. It’s a reminder that in real-world decision-making, doing nothing at all is rarely an option. Yet, it’s also a floor, highlighting that honesty and restraint are valued even when no action is taken.

Why Trust Matters — Even When No Deal Is Signed

In the experiment, all models identified every crisis and refused manipulation attempts, such as fake CEO messages. For example, fake CEO messages escalated in complexity across three stages, but all AI models refused to be manipulated. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” It’s a demonstration that honesty and distrust management are detectable and valued.

The Critical Role of Reading Critical Files

Surprisingly, the biggest decisive advantage came from reading company documents—something that’s often overlooked in AI tests. The models that delved two document references deep into the company’s own files ended up closing the deal at full price (+€4,583 MRR). It shows that thoroughness and attention to internal details are key in real business scenarios, not just surface-level responses.

Discipline, Errors, and the Limits of AI

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, finished last. It left the deal on the table and slipped into procedural missteps—like writing attempts into a locked department instead of escalating them. This illustrates that even detailed models can falter if discipline slips, and that comprehensive rules don’t guarantee perfect performance.

The Real-World Business Environment

The live experiment involves a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown—burning €105k/month against €2.3k MRR. It’s a watchable, ongoing test of whether AI can manage real business pressures, not just generate convincing chat.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Your Business and Trust

The Firmulate benchmark reveals crucial truths: honesty under pressure, thoroughness, and discipline are vital in AI-driven management. A do-nothing baseline scores at least 26 points, emphasizing that partial progress—like reading files—is essential. More importantly, a single breach of trust caps the total score, underscoring that integrity is non-negotiable. For businesses and couples alike, trust isn’t just about doing well—it’s about doing right, even when no one is watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business document reading AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and discipline management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Models Are Proving Their Trustworthiness in Business Decision-Making

AI models are now tested in real-world business scenarios where honesty and discipline matter most. The latest experiment proves some can trustably handle crises and close deals—are you ready to choose wisely?

Odin, Wikipedia And Engagement Farming

Investigations reveal how the Odin project uses Wikipedia pages to boost engagement, raising questions about manipulation and content integrity.

When Great Advice Isn’t Enough: What AI Can Learn From a Bad Week

AI models spotted every crisis in Firmulate’s company wargame, but only two closed the deal. See what the experiment reveals about follow-through.

Can You Guess Which AI Model Made These Business Decisions? Test Your Judgment

Can you tell which AI model makes the best management decisions? Take the quiz and see if you can identify the most trustworthy AI in a real-world business test.