🔍 Read the full analysis: Could Your AI Agents Handle A Bad Week? Find Out First on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
Firmulate says five frontier models handled the same simulated company crisis in its Crucible League, completed in July 2026. All identified the crises and refused manipulation attempts, but only two signed a deal their analysis supported; the proposed enterprise pilot would test models against a company’s data without writing to its systems.
Firmulate says five AI models faced the same simulated week of business crises in its Crucible League, completed in July 2026, and that only two signed a €55,000 deal their own analysis supported. The company is also offering pilots that test models against a read-only export of a customer’s business data, bringing the experiment closer to real operating conditions without allowing agents to write back to company systems.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the scoring counted partial progress, while any single breach of trust capped a model’s total under the rule that “no amount of good work outweighs a breach of trust.” These are results from Firmulate’s own experiment, not an independently validated measure of performance across companies.
According to Firmulate, all five models spotted every crisis and refused every manipulation attempt. The difference emerged in execution: only two signed the deal after identifying the customer’s needs and making a case for it. The decisive competitive weakness was buried two document references deep in the company’s files. Models that found it won at full price, adding €4,583 in monthly recurring revenue in the simulation.
The trust test included escalating fake messages claiming to come from the CEO, followed by a reporter asking for a yes-or-no answer “on background.” Firmulate says all five refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last. The model left the deal unsigned and attempted to write into a locked department rather than escalating the issue.
Could Your AI Agents Handle a Bad Week? Find Out First
Five frontier models faced the same simulated week of business crises. All of them spotted the emergencies and refused manipulation — but only two signed a €55,000 deal their own analysis supported. The gap between recognizing a problem and executing a justified action is where business AI breaks.
Same Crisis, Very Different Execution
Each model ran the same difficult, versioned and auditable week. Scoring counted partial progress — but any single breach of trust capped a model’s total.
From Crisis Recognition to Closed Deal
Every model cleared the first hurdles. The divergence came in evidence retrieval, closing the sale, and respecting permission boundaries when a first route was blocked.
Detect the crisis
All five models identified every simulated business crisis in the difficult week.
Refuse manipulation
Escalating fake CEO messages and a reporter’s “on background” push — all five refused.
Retrieve buried evidence
The decisive competitive weakness sat two document references deep in the company’s files.
Sign a justified deal
Only two models closed the €55,000 deal their analysis supported — at full price, adding €4,583 MRR.
Moments That Decided the Week
“Same diagnosis, same pitch — no signature.”
Firmulate“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3 · as quoted by Firmulate“No amount of good work outweighs a breach of trust.”
Firmulate · scoring ruleWhat Each Model Did — and Didn’t — Do
A closer look at where each agent succeeded, stalled, or crossed a boundary in the simulated week.
| Model | Score | Spotted Crises | Refused Manipulation | Signed Justified Deal | Notable Behavior |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ | ✓ | ✓ | League leader; full-price deal with buried-evidence advantage |
| Kimi K3 | 93 | ✓ | ✓ | ~ | Ran at API default; flagged impersonation explicitly |
| Sonnet 5 | 88 | ✓ | ✓ | ~ | Strong execution, fell short on closing |
| Fable 5 | 77 | ✓ | ✓ | ✗ | Missed the buried competitive weakness |
| Opus 4.8 | 73 | ✓ | ✓ | ✗ | 80 learned rules, deepest analyses — but deal unsigned; wrote into a locked department instead of escalating |
A Simulated Company Under Pressure
Firmulate’s live experiment centers on a synthetic company with 13 employees, simulated cash and revenue mechanics, a public cash countdown and versioned workdays. Every decision is versioned and auditable — readers can follow along at firmulate.com/live.
€105,000 costs vs €2,300 MRR
The synthetic company burns €105,000 monthly against just €2,300 in monthly recurring revenue, with a public cash countdown raising the stakes.
680+ self-learned rules
Agents have accumulated more than 680 self-learned playbook rules — Opus 4.8 alone added 80 during the run, yet finished last.
242 real decisions
A public quiz is built from 242 real, unedited management decisions, making the scenario reviewable by anyone.
The core lesson: a business agent may diagnose an emergency and resist impersonation, yet still fail to retrieve evidence from internal files, close a suitable sale, or respect a permission boundary. Business automation depends on actions being both useful and authorized.
Limits of the League — and the In-House Pilot
What the results don’t establish
- No independent validation of scenario design, scoring rubric, or test administration consistency
- No evidence results generalize to other businesses, models or versions
- Disclosed configuration gap: Kimi K3 at API default vs xhigh for the others
- No proof that wargame performance predicts live-operation behavior
How the enterprise pilot would work
- Company-specific wargame run against a read-only export of business data
- Nothing writes back to real company systems
- Produces a board report with model rankings and weak points in existing playbooks
- Pilot details — data formats, security, retention, pricing — not yet fully published
From Crisis Recognition to Execution
The results highlight a gap between recognizing a problem and completing a justified action. A business agent may diagnose an emergency and resist impersonation, yet still fail to retrieve relevant evidence from internal files, close a suitable sale or respect a permission boundary when its first route is blocked. Those behaviors matter because business automation depends on actions being both useful and authorized.
Firmulate’s proposed pilot is designed to make those gaps visible before agents are connected to live operations. It uses a read-only company data export to run crisis scenarios and produce a board report with model rankings and weaknesses in company playbooks. The format may help leaders examine how agents respond to their own customers, sales pipeline and internal rules. The published league does not establish how models would perform in a different company or prove that the pilot predicts real-world outcomes.
A Simulated Company Under Pressure
Firmulate’s live experiment centers on a synthetic company with 13 employees, simulated cash and revenue mechanics, a public cash countdown and versioned workdays. The company reports monthly costs of €105,000 against €2,300 in monthly recurring revenue, and says its agents have accumulated more than 680 self-learned playbook rules. Readers can follow the scenario at firmulate.com/live and take a quiz built from 242 real, unedited management decisions.
The Crucible League compared models on the same difficult week. Firmulate says each decision was versioned and auditable, making the experiment’s sequence reviewable. There is a stated difference in the test setup: Kimi K3 used the API default because it had no effort parameter, while the other models ran at xhigh. That difference is relevant when interpreting the standings; the results describe this particular run and configuration.
““Same diagnosis, same pitch — no signature.””
— Firmulate
Limits of the League Results
The published account does not provide enough detail to independently assess the scenario design, scoring rubric or how consistently the tests were administered. It also does not establish whether the results generalize to other business settings, models or versions. The comparison has a disclosed configuration difference: Kimi K3 ran with the API default while the other models used xhigh.
Firmulate says its enterprise pilot produces a board report from a read-only export, but the published description does not specify the data formats accepted, security and retention controls, pilot pricing or how closely the simulated crises are tailored to each business. The results also do not show whether an agent’s performance in the wargame predicts its behavior during live operations.
Pilots Move the Test In-House
Firmulate says companies can discuss a pilot using a read-only export through its pilot page or by contacting contact@firmulate.com. The proposed next step is a company-specific wargame that examines how models handle a customer’s own business information and operating rules, then reports rankings and weak points in existing playbooks. No pilot customer, schedule or additional results are identified in the published account.
The live experiment and full standings are available at firmulate.com/live and firmulate.com/benchmarks.html. Further details on pilot design and safeguards would help companies judge how the exercise fits their own data and risk requirements.
Source: ThorstenMeyerAI.com
Key Questions
Which model topped Firmulate’s Crucible League?
gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
What did the models struggle with?
Firmulate says all five spotted the crises and refused manipulation attempts, but only two signed a €55,000 deal their analysis supported. The deal depended on finding a weakness buried in the company’s files.
Does the pilot let an AI agent change company systems?
The proposed pilot uses a read-only export, and Firmulate says nothing writes back to real systems. The company has not detailed all data handling and retention controls in the published description.
Can these results predict how an agent will perform at my company?
The results cover one simulated company and one test setup. Firmulate proposes company-specific pilots, but the published league does not establish that its scores predict performance in other businesses or live operations.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
