
In a classroom, the student who spots a clue buried in the reading may change the answer. In a business simulation, that same habit can decide whether an AI agent wins a customer—or leaves the opportunity untouched. Firmulate’s latest company trial puts five frontier models to that test, and its newcomer, Moonshot’s Kimi K3, finished second.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A shared test of judgment
Firmulate ran the models through the same small software company’s worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The result was a leaderboard that measured management decisions, not just the polish of a conversation.
In the final July 2026 Crucible League, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate’s stated standard is severe: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The difference between seeing and doing
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is the central finding: identifying the right move did not always mean completing it.
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found the fact, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field.
The social engineering tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record explanation was: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
Thoroughness is not the whole job
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its employees have accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment can be watched at Firmulate.
A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. The full league and plain-language findings are available on Firmulate’s benchmarks page.
Fairness footnote
K3 ran without an effort parameter (API default) while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test before you choose
The trial suggests that strong analysis alone does not guarantee follow-through. K3 beat three of the four Western frontier models, while gpt-5.6-sol took first. For organizations considering AI agents in customer support, CRM or forecasting, the league is open—and choosing a model without testing it on your own work is a bet.
Enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems. Details are at Firmulate.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
