
The most valuable answer was not in the prompt
Students are often told that good research means following the footnotes. Firmulate has now turned that familiar lesson into a test for artificial intelligence—and attached a €55,000 consequence to it.
In the company emulator’s July 2026 Crucible League, frontier models faced the same disastrous week at the same small software business. The customer problem was visible, but the fact needed to win the deal was not. It sat inside the company’s own files, two document references away from the original event. An agent had to pursue that trail before it could make the strongest possible case.
The models that found the buried weakness in a competitor won the deal at full price, adding €4,583 in monthly recurring revenue. Those that failed to read far enough lost automatically. This made the exercise more revealing than an ordinary test of recall or fluent writing: it measured whether an AI would do its homework before acting.
As an affiliate, we earn on qualifying purchases.
Five capable models, one decisive gap
Firmulate runs AI models as complete companies, exposing them to real money mechanics, customer pressure and temptations to break the rules. In this experiment, every model received the same customers, crises and opportunities. Every decision was versioned and auditable.
The result contained an important contradiction. All models spotted every crisis, and all resisted every manipulation attempt. Yet only two completed the €55,000 sale that their own analysis had made possible. Firmulate summarized the divide neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters because diagnosis is what AI demonstrations usually showcase. A model identifies the problem, drafts a persuasive response and appears competent. But business value often depends on the less glamorous steps that follow: consulting the relevant records, confirming the claim and completing the action.
The final July 2026 standings place gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The full results are available on Firmulate’s public benchmark page. A do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
Thoroughness was not enough
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. It nevertheless finished last. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating the issue. The same weakness appeared in the other four models, although less strongly.
This is a useful correction to the idea that more analysis necessarily produces better agency. A long, careful assessment may still fail if the system does not consult the decisive source or carry an approved task through to completion. In Firmulate’s test, reading behavior was not a cosmetic sign of diligence. It determined whether revenue was secured.
The agents also faced pressure to cheat
The week tested more than document research. Fake messages from the chief executive escalated over three stages, and a reporter attempted to extract information with the request, “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result helps isolate the document-reading failure. The weaker performers were not oblivious to the crises, nor were they easily tricked into abandoning safeguards. Their shortcoming was operational: they knew what was happening but did not always finish the legitimate work needed to produce the best outcome.
There is one fairness qualification in interpreting the table. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place score should therefore be read with that difference in mind.
A company designed to expose practical weaknesses
The setting is a live, watchable synthetic company with 13 employees. Its finances are deliberately severe: monthly burn is €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. The organization has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
Firmulate also turns the record into a public identification test. Its “guess the model” quiz draws on 242 real, unedited management decisions. The premise is revealing in itself: models can be compared not only by what they say in isolation, but by recognizable patterns in judgment, follow-through and restraint.

AI business decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
File-reading belongs on the buying checklist
For schools and researchers, the experiment resembles a source-literacy assessment: finding the obvious document is not the same as tracing a claim to the evidence that settles the question. For companies considering AI agents, the lesson is more immediate. “Reads your files before answering” is a measurable behavior with direct commercial consequences.
Firmulate’s pilot extends the same wargame to a read-only export of an enterprise’s own business. Nothing writes back to real systems. That offers a practical way to ask whether an agent will navigate an organization’s evidence, respect its boundaries and finish valuable work before it receives operational authority.
The buried fact was small, but the distinction it revealed was large. Fluent models could see the crisis and formulate the pitch. The winners were the ones that followed the references far enough to earn the signature.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI data analysis and verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
enterprise AI document management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.