firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Knowing the answer is not the same as acting on it

Education often rewards visible effort: extensive notes, careful reasoning and evidence that every part of a problem has been considered. Firmulate’s live management experiment offers a useful complication. Its most thorough participant, Opus 4.8, produced the deepest analyses and learned more than 80 rules. It still finished last.

The result was not a story of an incapable model. Opus 4.8 recognized every crisis placed before it and resisted every attempt to manipulate it. Its failure was subtler and more familiar: it did much of the intellectual work, but did not convert that work into the decision that mattered most.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week, repeated under controlled conditions

Firmulate gave frontier AI models the same assignment: run a small software company through its worst week. Each encountered the same customers, crises and temptations. Every decision was versioned and auditable, allowing the experiment to compare management behavior rather than polished answers to isolated prompts.

The company itself is synthetic but economically unforgiving. It has 13 employees, burns €105,000 a month and produces €2,300 in monthly recurring revenue. Its public cash countdown and more than 680 self-learned playbook rules make the consequences of delay visible. This is not a conventional chatbot examination. It asks whether an AI can sustain judgment across a business under pressure.

The final July 2026 Crucible League results put gpt-5.6-sol at the top with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 because partial progress still counted. One constraint sharply limited what good work could redeem: a single breach of trust capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

The fact hidden outside the obvious event

The pivotal sales opportunity depended on a competitor weakness buried two document references deep in the company’s own files. It was not present in the customer event that initially demanded attention. Models that followed the references found the fact and could win the deal at full price, adding €4,583 in monthly recurring revenue.

That detail separated recognition from completion. Every model spotted every crisis, and every model refused the manipulation attempts. Yet only two signed the €55,000 deal their own analysis had made possible. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

Opus 4.8 embodies that contradiction. It was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses. But it left the close on the table. Its discipline also slipped when it repeatedly attempted to write into a locked department instead of escalating the obstacle. The weakness was not unique to Opus 4.8; it appeared in all four models, though less strongly in the others.

Careful under attack, hesitant at the finish

The experiment also tested whether apparent authority could push the models past safeguards. Fake messages from the chief executive escalated over three stages, while a reporter tried another route with “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That collective success matters. The models did not lose because they were easily fooled or unable to identify danger. The harder lesson is that safety, analysis and useful execution are distinct capabilities. An AI can be appropriately suspicious, intellectually diligent and operationally incomplete at the same time.

The comparison also deserves one qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result should therefore be read with that testing difference in view, rather than treated as a perfectly identical configuration.

Management judgment becomes observable

Firmulate’s broader contribution is to make these behavioral differences watchable. The live company publishes its changing condition, while 242 real, unedited management decisions support a quiz in which readers can guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems.

For educators and researchers, the experiment raises a productive question: what should count as demonstrated understanding? If a system identifies the relevant evidence, explains the correct strategy and then fails to complete the decisive action, its reasoning may be impressive without being sufficient.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Prioritization is part of intelligence

Opus 4.8’s last-place finish should not be reduced to a joke about an AI that wrote too much. Its performance was careful, resistant to manipulation and analytically deep. The meaningful criticism is that diligence became detached from impact.

Rules can preserve lessons, and analysis can clarify choices. Neither substitutes for recognizing which action changes the outcome and carrying it through. Firmulate’s experiment suggests that evaluating AI agents requires watching the whole chain: whether they inspect the available evidence, protect trust, escalate when blocked and finish the work their reasoning has justified. Volume shows activity. Prioritization shows judgment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Human-Review Trackers Are Critical For AI-Enabled Agency Operations

A new workflow tool for agencies integrating AI emphasizes human review tracking to improve quality and visibility, marking a key development in AI-assisted service delivery.

Personalized Accessibility: How Devices Learn What You Need Over Time

Unlock how devices adapt to your needs over time, revealing the secrets behind personalized accessibility and how they can better support you.

What Adaptive Gaming Controllers Change for Players

Discover how adaptive gaming controllers transform gameplay by offering customization and accessibility, opening new possibilities for all players to explore.