firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Can an AI spot a crisis, resist a convincing lie and still fail to do the job? Firmulate’s live experiment turns that question into a practical lesson in management: knowing what should happen and following through are not the same thing. For readers who care about how people learn and make decisions, the test offers a watchable case study in judgment under pressure.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The experiment puts frontier AI models in charge of the same small software company during its worst week. They face the same customers, crises and temptations. Firmulate records each decision so viewers can follow what happened and why.

Recognition is only the first test

In the final Crucible League, published in July 2026, gpt-5.6-sol ranked first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark gives partial credit for progress, but a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The headline finding was strikingly consistent: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” The gap between understanding a good decision and carrying it through is hard to see in a chat demo. Here, it became part of the story.

The clue was already in the company’s files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The episode makes a useful point about business judgment: success may depend on connecting information that is available but easy to overlook.

The experiment also tested social engineering. Fake messages from a CEO escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

More effort did not guarantee a better finish

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also test their intuition against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. The live company has 13 synthetic employees, real money mechanics, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Its burn is €105k/month against €2.3k MRR. The experiment is real and watchable at firmulate.com.

From watching to a company-specific pilot

For business leaders, the next question is whether the same kind of test can reveal weaknesses in their own company’s playbooks. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that representation. The resulting board report includes model rankings and weak points in the company’s own responses. Nothing writes back to real systems.

That makes the pilot a move from observing a public experiment to examining your own operating assumptions. It can show how AI agents handle pressure in the context of your customers, processes and rules, before anyone considers giving them a role in live systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s experiment shows that crisis recognition and refusal are not enough: models also have to find relevant information, follow through and respect boundaries. Enterprises can explore those behaviors against a read-only export of their own business, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Dyslexia Support Tech: How Reading Tools Reduce Visual Stress

Supporting dyslexic readers with customizable tools can significantly reduce visual stress, but discovering how these features work may change your experience forever.

Revealing The AI Opportunities Zero-Sum Enthusiasts Fail To See

Analysis of Thorsten Meyer’s interview with Eric Vishria highlights misconceptions in AI market competition and emphasizes the importance of differentiation and hardware control.

Best AI Automation Solutions For Small Businesses This Year

Discover the best AI automation tools for small businesses this year, featuring 15 top options to streamline operations, marketing, sales, and more.

When the Most Diligent AI Still Fails the Test

Firmulate’s most diligent AI found every crisis and learned 80 rules, yet finished last—a lesson in why analysis without decisive action falls short.