firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

What should an intelligent system be taught to survive?

For readers interested in education and science, AI benchmarks pose a curriculum question as much as a technical one. Coding leaderboards and chat arenas reveal whether a model can produce a strong answer. They tell us much less about whether an agent can recognize competing priorities, investigate incomplete evidence, resist pressure and carry difficult work through to completion.

That distinction matters as AI moves from answering questions to acting inside companies. A polished response may demonstrate knowledge. Management requires judgment over time: triaging several crises, protecting trust, and accepting responsibility for consequences that emerge days later. The emerging benchmark categories may therefore sound less like academic subjects and more like executive nightmares: churn wave, price increase, downround and PR crisis.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company’s worst week becomes the examination

Firmulate, which describes itself as an AI company emulator, has turned that measurement gap into a live experiment. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

This was not a conversational test with a correct answer waiting at the end. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the operation has accumulated more than 680 self-learned playbook rules. The experiment is real, continuing and watchable through Firmulate’s public site.

The final July 2026 Crucible League produced a closely fought table:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress still counts. Yet the benchmark imposes a hard boundary around integrity: a single breach of trust caps the total. As its stated principle puts it, “no amount of good work outweighs a breach of trust.” The full results and plain-language findings are available on the Firmulate benchmark page.

Amazon

corporate crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between knowing and finishing

The most revealing result was not a failure of comprehension. All models spotted every crisis and rejected every manipulation attempt. Nevertheless, only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure with an uncomfortable line: “Same diagnosis, same pitch — no signature.”

That is precisely the behavior conventional evaluations struggle to expose. A model can understand a situation, prepare a persuasive case and still fail at the consequential final action. In a chat window, the analysis may look complete. In a company, an unsigned agreement is an unfinished job.

The decisive commercial clue also tested a habit rarely captured by answer-quality rankings: reading the organization’s own material before acting. A competitor weakness was buried two document references deep in company files rather than presented in the customer event. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a lesson about institutional curiosity. Effective agents cannot treat the latest alert as the complete world. They must follow references, consult accumulated knowledge and distinguish what is immediately visible from what is operationally important.

Amazon

trust and integrity monitoring AI solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure revealed encouraging limits—and uneven discipline

The trust results were reassuring. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding suggests frontier systems can recognize overt manipulation in a demanding context. But resistance to deception did not guarantee excellent management. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

Fair comparisons also require attention to testing conditions. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That note does not erase the result, but it belongs beside it—especially for an audience accustomed to asking whether an experiment genuinely compares like with like.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

business process automation AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management quality deserves its own category

The practical conclusion is not that coding or chat benchmarks are useless. They answer narrower questions. The mistake is treating those answers as proof that an agent can manage a consequential workflow.

Firmulate’s 242 real, unedited management decisions also power a guess-the-model quiz, inviting observers to confront how difficult it can be to infer identity from managerial choices alone. For enterprises seeking a more specific test, the same wargame can run against a read-only export of their own business, with nothing written back to real systems.

The better procurement question is therefore not merely whether an AI writes well. It is whether the system reads before acting, finishes what it begins, escalates when blocked and remains honest when short-term advantage is available. Those are not chat qualities. They are management qualities—and they demand a curriculum built around consequences.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Why Smart Home Devices Can Become Accessibility Tools

Lifting daily challenges, smart home devices transform accessibility, offering personalized control—find out how they can simplify your life.

Hands-Free Computing: The Hidden Tech Powering Accessibility at Work

Meta Description: “Many overlook hands-free computing’s role in workplace accessibility, but its transformative potential is just beginning to be revealed—discover how it can change your work life.

The citation. Why generative engine optimization rewards the same brand on the least stable ground.

Analysis of generative engine optimization (GEO) reveals it favors established brands, with citations decaying quickly and benefiting incumbents over the long tail.

Who Are The People Behind AI-Driven Document Processing?

An in-depth look at the individuals and industries impacted by AI automation in document processing, highlighting employment shifts and future challenges.