AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine trusting an AI to handle your most sensitive business decisions—only to find it refuses to sign a simple deal or read critical files. For couples and companies alike, trust is everything. But how do we measure it?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark That Tests Trust and Discipline

Recently, a groundbreaking experiment called the Firmulate AI benchmark put four advanced AI models through a simulated week of running a small software company. This wasn’t about chatty responses or clever tricks; it was about managing crises, reading critical documents, and making honest decisions under pressure.

The Do-Nothing Baseline: Why It Scores 26 Points

One striking finding emerged: even a ‘do-nothing’ AI, which merely observes and refrains from acting, scores 26 out of 100 points. This score exists because partial progress—like reading a file or recognizing a crisis—counts towards the total. It’s a reminder that in real-world decision-making, doing nothing at all is rarely an option. Yet, it’s also a floor, highlighting that honesty and restraint are valued even when no action is taken.

Why Trust Matters — Even When No Deal Is Signed

In the experiment, all models identified every crisis and refused manipulation attempts, such as fake CEO messages. For example, fake CEO messages escalated in complexity across three stages, but all AI models refused to be manipulated. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” It’s a demonstration that honesty and distrust management are detectable and valued.

The Critical Role of Reading Critical Files

Surprisingly, the biggest decisive advantage came from reading company documents—something that’s often overlooked in AI tests. The models that delved two document references deep into the company’s own files ended up closing the deal at full price (+€4,583 MRR). It shows that thoroughness and attention to internal details are key in real business scenarios, not just surface-level responses.

Discipline, Errors, and the Limits of AI

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, finished last. It left the deal on the table and slipped into procedural missteps—like writing attempts into a locked department instead of escalating them. This illustrates that even detailed models can falter if discipline slips, and that comprehensive rules don’t guarantee perfect performance.

The Real-World Business Environment

The live experiment involves a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown—burning €105k/month against €2.3k MRR. It’s a watchable, ongoing test of whether AI can manage real business pressures, not just generate convincing chat.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Your Business and Trust

The Firmulate benchmark reveals crucial truths: honesty under pressure, thoroughness, and discipline are vital in AI-driven management. A do-nothing baseline scores at least 26 points, emphasizing that partial progress—like reading files—is essential. More importantly, a single breach of trust caps the total score, underscoring that integrity is non-negotiable. For businesses and couples alike, trust isn’t just about doing well—it’s about doing right, even when no one is watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business document reading AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and discipline management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Integrity Shines Under Pressure: A Real-World Test of Trust and Security

Real-world AI tests show models refusing manipulation in simulated crises, demonstrating integrity and trustworthiness essential for deploying AI in critical roles.

BYD Taking Responsibility Increases God’s Eye Use & Makes Vehicles Safer

BYD’s new city driving guarantee increases God’s Eye adoption, reduces accidents, and lowers insurance costs, marking a significant shift in EV safety and liability.

Google Will Soon Let You Tell It What Stories You Want To See In Your Discovery Feed

Google will soon enable users to specify story preferences in Discovery using natural language prompts, enhancing personalized news experiences.

Second project portfolio should take Hawaii university 100% solar-powered

Brigham Young University-Hawaii announces second phase of solar project to fully power campus and nearby facilities, including energy storage systems.