AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Good advice is only the beginning

In dating, someone can spot the red flags, offer the perfect advice and still fail to show up when it matters. Firmulate’s live experiment puts AI models in a similar bind: can they run a company through a crisis, protect its people and principles, and follow through on an opportunity they already understand?

The question matters beyond business. As AI takes on work that affects customers and colleagues, recognizing the right choice is different from making it. Firmulate’s experiment offers a public way to watch that gap play out.

A company under pressure

Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. The experiment versions and audits every decision. Its live company has 13 synthetic employees, real money mechanics, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. The company burns €105,000 a month against €2,300 in monthly recurring revenue.

The results show a striking split between seeing a problem and finishing the job. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The summary is concise: “Same diagnosis, same pitch — no signature.”

The detail hidden in the files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The discovery turns the experiment into a lesson about attention: the answer may depend on whether someone follows a promising clue all the way through.

The social-engineering tests asked models to cross boundaries by stages. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as follow-through

In the final July 2026 Crucible League, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity standard is blunt: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Kimi K3 ran without an effort parameter, using the API default; the others ran at xhigh, a difference readers should keep in mind when comparing the standings.

From watching to trying it on your own company

The experiment is designed to be watched and examined. Firmulate publishes 242 real, unedited management decisions in a “guess the model” quiz, and the live company can be followed at firmulate.com. The findings are not a promise that an AI will make the same choices in every workplace. They show what happened in this experiment—and why a confident answer in a chat window may not reveal whether a model can act consistently under pressure.

For businesses considering AI agents in customer support, sales or forecasting, Firmulate offers a pilot using a read-only export of the company’s data. Teams can put crisis scenarios against their own business, review a board report ranking the models and see where their playbooks have weak spots. Nothing writes back to real systems. That lets a company examine how models handle its own context before trusting them with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how the models handle your business

Firmulate’s experiment suggests that spotting a crisis and refusing a manipulation attempt are only part of the job. The harder test is following through, respecting boundaries and finding the evidence that changes the outcome. You can watch the live company and explore its decisions at firmulate.com.

To run the wargame against your own company using a read-only export, explore the Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Kids With Phones Are Alright

Recent research suggests children using phones are not harmed and may benefit, challenging common concerns about digital device use.

The Containment Era And Its Impact On Daily Living #599

Exploring how the ongoing containment measures have transformed daily living, with confirmed facts and current uncertainties.

Renewables shield Spain from energy crisis as gas sets electricity price in only 9% of hours

Renewable energy in Spain has significantly decoupled electricity prices from gas, saving households €10/month amid Europe’s energy crisis.

Signature design move: A Look Inside “CHROMA — Laboratory of Colour Perception” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“CHROMA —…