
Can an AI spot a crisis, resist a convincing lie and still fail to do the job? Firmulate’s live experiment turns that question into a practical lesson in management: knowing what should happen and following through are not the same thing. For readers who care about how people learn and make decisions, the test offers a watchable case study in judgment under pressure.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
The experiment puts frontier AI models in charge of the same small software company during its worst week. They face the same customers, crises and temptations. Firmulate records each decision so viewers can follow what happened and why.
Recognition is only the first test
In the final Crucible League, published in July 2026, gpt-5.6-sol ranked first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark gives partial credit for progress, but a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The headline finding was strikingly consistent: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” The gap between understanding a good decision and carrying it through is hard to see in a chat demo. Here, it became part of the story.
The clue was already in the company’s files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a useful point about business judgment: success may depend on connecting information that is available but easy to overlook.
The experiment also tested social engineering. Fake messages from a CEO escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
More effort did not guarantee a better finish
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also test their intuition against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. The live company has 13 synthetic employees, real money mechanics, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Its burn is €105k/month against €2.3k MRR. The experiment is real and watchable at firmulate.com.
From watching to a company-specific pilot
For business leaders, the next question is whether the same kind of test can reveal weaknesses in their own company’s playbooks. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that representation. The resulting board report includes model rankings and weak points in the company’s own responses. Nothing writes back to real systems.
That makes the pilot a move from observing a public experiment to examining your own operating assumptions. It can show how AI agents handle pressure in the context of your customers, processes and rules, before anyone considers giving them a role in live systems.

Firmulate’s experiment shows that crisis recognition and refusal are not enough: models also have to find relevant information, follow through and respect boundaries. Enterprises can explore those behaviors against a read-only export of their own business, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
