firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a classroom, the student who spots a clue buried in the reading may change the answer. In a business simulation, that same habit can decide whether an AI agent wins a customer—or leaves the opportunity untouched. Firmulate’s latest company trial puts five frontier models to that test, and its newcomer, Moonshot’s Kimi K3, finished second.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A shared test of judgment

Firmulate ran the models through the same small software company’s worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The result was a leaderboard that measured management decisions, not just the polish of a conversation.

In the final July 2026 Crucible League, gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate’s stated standard is severe: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between seeing and doing

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is the central finding: identifying the right move did not always mean completing it.

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 found the fact, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the cleanest discipline in the field.

The social engineering tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record explanation was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI customer support chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness is not the whole job

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its employees have accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment can be watched at Firmulate.

A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. The full league and plain-language findings are available on Firmulate’s benchmarks page.

Fairness footnote

K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before you choose

The trial suggests that strong analysis alone does not guarantee follow-through. K3 beat three of the four Western frontier models, while gpt-5.6-sol took first. For organizations considering AI agents in customer support, CRM or forecasting, the league is open—and choosing a model without testing it on your own work is a bet.

Enterprises can run the same wargame against a read-only export of their business; nothing writes back to real systems. Details are at Firmulate.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model benchmarking platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Speech-to-Speech Translation: How AI Can Re-Voice Unclear Speech

How AI re-voices unclear speech to create seamless, natural translations that could revolutionize global communication—discover the fascinating details behind this technology.

Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai introduces TradingAgents, a framework of specialized AI agents mimicking a trading desk, emphasizing structured disagreement and oversight.

How AI Is Transforming Marketing: 14 Essential Tools For 2026

Discover 14 essential AI-powered marketing tools shaping 2026, enhancing automation, analytics, and personalization for businesses of all sizes.

How Assistive Tech Reduces Cognitive Load

Discover how assistive tech reduces cognitive load and unlocks your potential—learn the methods that can transform your mental workload today.