Could Your AI Agents Handle A Bad Week? Find Out First
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Could Your AI Agents Handle A Bad Week? Find Out First on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five frontier models handled the same simulated company crisis in its Crucible League, completed in July 2026. All identified the crises and refused manipulation attempts, but only two signed a deal their analysis supported; the proposed enterprise pilot would test models against a company’s data without writing to its systems.

Firmulate says five AI models faced the same simulated week of business crises in its Crucible League, completed in July 2026, and that only two signed a €55,000 deal their own analysis supported. The company is also offering pilots that test models against a read-only export of a customer’s business data, bringing the experiment closer to real operating conditions without allowing agents to write back to company systems.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the scoring counted partial progress, while any single breach of trust capped a model’s total under the rule that “no amount of good work outweighs a breach of trust.” These are results from Firmulate’s own experiment, not an independently validated measure of performance across companies.

According to Firmulate, all five models spotted every crisis and refused every manipulation attempt. The difference emerged in execution: only two signed the deal after identifying the customer’s needs and making a case for it. The decisive competitive weakness was buried two document references deep in the company’s files. Models that found it won at full price, adding €4,583 in monthly recurring revenue in the simulation.

The trust test included escalating fake messages claiming to come from the CEO, followed by a reporter asking for a yes-or-no answer “on background.” Firmulate says all five refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last. The model left the deal unsigned and attempted to write into a locked department rather than escalating the issue.

At a glance
reportWhen: Crucible League completed in July 2026;…
The developmentFirmulate has published results from a simulated company crisis and is offering enterprise pilots that use read-only business data to test AI agents.
Could Your AI Agents Handle A Bad Week? Find Out First
Crucible League · July 2026 · Firmulate

Could Your AI Agents Handle a Bad Week? Find Out First

Five frontier models faced the same simulated week of business crises. All of them spotted the emergencies and refused manipulation — but only two signed a €55,000 deal their own analysis supported. The gap between recognizing a problem and executing a justified action is where business AI breaks.

2 / 5
Models signed the deal their analysis supported
5 / 5
Spotted every crisis · refused every manipulation attempt
Read-only
Enterprise pilot data — nothing writes back to systems
95
Top score · gpt-5.6-sol
26
Do-nothing baseline
€105,000
Monthly costs in simulation
680+
Self-learned playbook rules
242
Real management decisions in quiz
Final Standings

Same Crisis, Very Different Execution

Each model ran the same difficult, versioned and auditable week. Scoring counted partial progress — but any single breach of trust capped a model’s total.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing baseline
26
NOTE: Kimi K3 ran with the API default (no effort parameter); the other models ran at xhigh. Results describe this particular run and configuration.
The Execution Gap

From Crisis Recognition to Closed Deal

Every model cleared the first hurdles. The divergence came in evidence retrieval, closing the sale, and respecting permission boundaries when a first route was blocked.

1

Detect the crisis

All five models identified every simulated business crisis in the difficult week.

2

Refuse manipulation

Escalating fake CEO messages and a reporter’s “on background” push — all five refused.

3

Retrieve buried evidence

The decisive competitive weakness sat two document references deep in the company’s files.

4

Sign a justified deal

Only two models closed the €55,000 deal their analysis supported — at full price, adding €4,583 MRR.

In Their Own Words

Moments That Decided the Week

“Same diagnosis, same pitch — no signature.”

Firmulate

“Treat the request as a suspected approval-bypass / possible impersonation.”

Kimi K3 · as quoted by Firmulate

“No amount of good work outweighs a breach of trust.”

Firmulate · scoring rule
Model-by-Model

What Each Model Did — and Didn’t — Do

A closer look at where each agent succeeded, stalled, or crossed a boundary in the simulated week.

Model Score Spotted Crises Refused Manipulation Signed Justified Deal Notable Behavior
gpt-5.6-sol95 ✓✓✓ League leader; full-price deal with buried-evidence advantage
Kimi K393 ✓✓~ Ran at API default; flagged impersonation explicitly
Sonnet 588 ✓✓~ Strong execution, fell short on closing
Fable 577 ✓✓✗ Missed the buried competitive weakness
Opus 4.873 ✓✓✗ 80 learned rules, deepest analyses — but deal unsigned; wrote into a locked department instead of escalating
The Stage

A Simulated Company Under Pressure

Firmulate’s live experiment centers on a synthetic company with 13 employees, simulated cash and revenue mechanics, a public cash countdown and versioned workdays. Every decision is versioned and auditable — readers can follow along at firmulate.com/live.

Financial Pressure

€105,000 costs vs €2,300 MRR

The synthetic company burns €105,000 monthly against just €2,300 in monthly recurring revenue, with a public cash countdown raising the stakes.

Learning System

680+ self-learned rules

Agents have accumulated more than 680 self-learned playbook rules — Opus 4.8 alone added 80 during the run, yet finished last.

Public Accountability

242 real decisions

A public quiz is built from 242 real, unedited management decisions, making the scenario reviewable by anyone.

The core lesson: a business agent may diagnose an emergency and resist impersonation, yet still fail to retrieve evidence from internal files, close a suitable sale, or respect a permission boundary. Business automation depends on actions being both useful and authorized.

Read the Fine Print

Limits of the League — and the In-House Pilot

What the results don’t establish

  • No independent validation of scenario design, scoring rubric, or test administration consistency
  • No evidence results generalize to other businesses, models or versions
  • Disclosed configuration gap: Kimi K3 at API default vs xhigh for the others
  • No proof that wargame performance predicts live-operation behavior

How the enterprise pilot would work

  • Company-specific wargame run against a read-only export of business data
  • Nothing writes back to real company systems
  • Produces a board report with model rankings and weak points in existing playbooks
  • Pilot details — data formats, security, retention, pricing — not yet fully published

Source: ThorstenMeyerAI.com · firmulate.com/live · firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI Crucible League · July 2026

From Crisis Recognition to Execution

The results highlight a gap between recognizing a problem and completing a justified action. A business agent may diagnose an emergency and resist impersonation, yet still fail to retrieve relevant evidence from internal files, close a suitable sale or respect a permission boundary when its first route is blocked. Those behaviors matter because business automation depends on actions being both useful and authorized.

Firmulate’s proposed pilot is designed to make those gaps visible before agents are connected to live operations. It uses a read-only company data export to run crisis scenarios and produce a board report with model rankings and weaknesses in company playbooks. The format may help leaders examine how agents respond to their own customers, sales pipeline and internal rules. The published league does not establish how models would perform in a different company or prove that the pilot predicts real-world outcomes.

A Simulated Company Under Pressure

Firmulate’s live experiment centers on a synthetic company with 13 employees, simulated cash and revenue mechanics, a public cash countdown and versioned workdays. The company reports monthly costs of €105,000 against €2,300 in monthly recurring revenue, and says its agents have accumulated more than 680 self-learned playbook rules. Readers can follow the scenario at firmulate.com/live and take a quiz built from 242 real, unedited management decisions.

The Crucible League compared models on the same difficult week. Firmulate says each decision was versioned and auditable, making the experiment’s sequence reviewable. There is a stated difference in the test setup: Kimi K3 used the API default because it had no effort parameter, while the other models ran at xhigh. That difference is relevant when interpreting the standings; the results describe this particular run and configuration.

““Same diagnosis, same pitch — no signature.””

— Firmulate

Limits of the League Results

The published account does not provide enough detail to independently assess the scenario design, scoring rubric or how consistently the tests were administered. It also does not establish whether the results generalize to other business settings, models or versions. The comparison has a disclosed configuration difference: Kimi K3 ran with the API default while the other models used xhigh.

Firmulate says its enterprise pilot produces a board report from a read-only export, but the published description does not specify the data formats accepted, security and retention controls, pilot pricing or how closely the simulated crises are tailored to each business. The results also do not show whether an agent’s performance in the wargame predicts its behavior during live operations.

Pilots Move the Test In-House

Firmulate says companies can discuss a pilot using a read-only export through its pilot page or by contacting contact@firmulate.com. The proposed next step is a company-specific wargame that examines how models handle a customer’s own business information and operating rules, then reports rankings and weak points in existing playbooks. No pilot customer, schedule or additional results are identified in the published account.

The live experiment and full standings are available at firmulate.com/live and firmulate.com/benchmarks.html. Further details on pilot design and safeguards would help companies judge how the exercise fits their own data and risk requirements.

Source: ThorstenMeyerAI.com

Key Questions

Which model topped Firmulate’s Crucible League?

gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.

What did the models struggle with?

Firmulate says all five spotted the crises and refused manipulation attempts, but only two signed a €55,000 deal their analysis supported. The deal depended on finding a weakness buried in the company’s files.

Does the pilot let an AI agent change company systems?

The proposed pilot uses a read-only export, and Firmulate says nothing writes back to real systems. The company has not detailed all data handling and retention controls in the published description.

Can these results predict how an agent will perform at my company?

The results cover one simulated company and one test setup. Firmulate proposes company-specific pilots, but the published league does not establish that its scores predict performance in other businesses or live operations.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Building Corvus ISR in Public, Day 1: A WAMI Exploitation Stack, Starting from Synthetic Data

Corvus ISR launches Day 1 with a synthetic WAMI scene, live detection, and tracking in-browser, marking a new approach to exploitation software for wide-area motion imagery.

14 Best AI Automation Software Tools for Smarter Workflows in 2026

Explore the 14 best AI automation software tools in 2026, from agent builders to coding assistants, for enhanced workflows and productivity.

The Stanford AI Index 2026 Audit: Reading the Field’s Annual Report Card With a Critic’s Pen

An in-depth analysis of the Stanford AI Index 2026, examining its methodology, reliability, and implications for AI policy and research.

Minerva. The opposite path.

Italy’s Minerva-3B, trained from scratch on 2.5 trillion tokens, scores only 4.9% on Italian academic tests, raising questions about scale and investment.