
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Good advice is only the beginning
In dating, someone can spot the red flags, offer the perfect advice and still fail to show up when it matters. Firmulate’s live experiment puts AI models in a similar bind: can they run a company through a crisis, protect its people and principles, and follow through on an opportunity they already understand?
The question matters beyond business. As AI takes on work that affects customers and colleagues, recognizing the right choice is different from making it. Firmulate’s experiment offers a public way to watch that gap play out.
A company under pressure
Firmulate ran frontier AI models through the same small software company’s worst week: the same customers, crises and temptations. The experiment versions and audits every decision. Its live company has 13 synthetic employees, real money mechanics, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. The company burns €105,000 a month against €2,300 in monthly recurring revenue.
The results show a striking split between seeing a problem and finishing the job. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The summary is concise: “Same diagnosis, same pitch — no signature.”
The detail hidden in the files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The discovery turns the experiment into a lesson about attention: the answer may depend on whether someone follows a promising clue all the way through.
The social-engineering tests asked models to cross boundaries by stages. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as follow-through
In the final July 2026 Crucible League, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity standard is blunt: “no amount of good work outweighs a breach of trust.”
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models. Kimi K3 ran without an effort parameter, using the API default; the others ran at xhigh, a difference readers should keep in mind when comparing the standings.
From watching to trying it on your own company
The experiment is designed to be watched and examined. Firmulate publishes 242 real, unedited management decisions in a “guess the model” quiz, and the live company can be followed at firmulate.com. The findings are not a promise that an AI will make the same choices in every workplace. They show what happened in this experiment—and why a confident answer in a chat window may not reveal whether a model can act consistently under pressure.
For businesses considering AI agents in customer support, sales or forecasting, Firmulate offers a pilot using a read-only export of the company’s data. Teams can put crisis scenarios against their own business, review a board report ranking the models and see where their playbooks have weak spots. Nothing writes back to real systems. That lets a company examine how models handle its own context before trusting them with live work.

See how the models handle your business
Firmulate’s experiment suggests that spotting a crisis and refusing a manipulation attempt are only part of the job. The harder test is following through, respecting boundaries and finding the evidence that changes the outcome. You can watch the live company and explore its decisions at firmulate.com.
To run the wargame against your own company using a read-only export, explore the Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
