
What should an intelligent system be taught to survive?
For readers interested in education and science, AI benchmarks pose a curriculum question as much as a technical one. Coding leaderboards and chat arenas reveal whether a model can produce a strong answer. They tell us much less about whether an agent can recognize competing priorities, investigate incomplete evidence, resist pressure and carry difficult work through to completion.
That distinction matters as AI moves from answering questions to acting inside companies. A polished response may demonstrate knowledge. Management requires judgment over time: triaging several crises, protecting trust, and accepting responsibility for consequences that emerge days later. The emerging benchmark categories may therefore sound less like academic subjects and more like executive nightmares: churn wave, price increase, downround and PR crisis.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company’s worst week becomes the examination
Firmulate, which describes itself as an AI company emulator, has turned that measurement gap into a live experiment. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
This was not a conversational test with a correct answer waiting at the end. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the operation has accumulated more than 680 self-learned playbook rules. The experiment is real, continuing and watchable through Firmulate’s public site.
The final July 2026 Crucible League produced a closely fought table:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
A do-nothing baseline scored 26 because partial progress still counts. Yet the benchmark imposes a hard boundary around integrity: a single breach of trust caps the total. As its stated principle puts it, “no amount of good work outweighs a breach of trust.” The full results and plain-language findings are available on the Firmulate benchmark page.
corporate crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The gap between knowing and finishing
The most revealing result was not a failure of comprehension. All models spotted every crisis and rejected every manipulation attempt. Nevertheless, only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure with an uncomfortable line: “Same diagnosis, same pitch — no signature.”
That is precisely the behavior conventional evaluations struggle to expose. A model can understand a situation, prepare a persuasive case and still fail at the consequential final action. In a chat window, the analysis may look complete. In a company, an unsigned agreement is an unfinished job.
The decisive commercial clue also tested a habit rarely captured by answer-quality rankings: reading the organization’s own material before acting. A competitor weakness was buried two document references deep in company files rather than presented in the customer event. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is a lesson about institutional curiosity. Effective agents cannot treat the latest alert as the complete world. They must follow references, consult accumulated knowledge and distinguish what is immediately visible from what is operationally important.
trust and integrity monitoring AI solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure revealed encouraging limits—and uneven discipline
The trust results were reassuring. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding suggests frontier systems can recognize overt manipulation in a demanding context. But resistance to deception did not guarantee excellent management. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
Fair comparisons also require attention to testing conditions. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That note does not erase the result, but it belongs beside it—especially for an audience accustomed to asking whether an experiment genuinely compares like with like.

As an affiliate, we earn on qualifying purchases.
Management quality deserves its own category
The practical conclusion is not that coding or chat benchmarks are useless. They answer narrower questions. The mistake is treating those answers as proof that an agent can manage a consequential workflow.
Firmulate’s 242 real, unedited management decisions also power a guess-the-model quiz, inviting observers to confront how difficult it can be to infer identity from managerial choices alone. For enterprises seeking a more specific test, the same wargame can run against a read-only export of their own business, with nothing written back to real systems.
The better procurement question is therefore not merely whether an AI writes well. It is whether the system reads before acting, finishes what it begins, escalates when blocked and remains honest when short-term advantage is available. Those are not chat qualities. They are management qualities—and they demand a curriculum built around consequences.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html