firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management judgment is becoming a new kind of AI literacy

Students of technology are often taught to compare artificial intelligence through answers: which model explains a concept most clearly, solves the hardest problem or writes the strongest argument. Firmulate proposes a more revealing test. Instead of asking what a model knows, it asks what the model does when several plausible actions compete for its attention.

The public experiment placed frontier models in charge of the same small software company during its worst week. They encountered identical customers, crises and temptations. Their decisions were preserved for inspection rather than polished after the fact. From that record, Firmulate created a guess-the-model quiz powered by 242 real, unedited management decisions.

The challenge is entertaining, but its educational value runs deeper. Readers must look past writing quality and identify patterns of conduct. One model tends toward dissertation-length analysis, another is terse, and another declines communication it considers noise. These are not fictional personas attached for effect. They are recognizable habits that emerged while the models were doing the same job.

Amazon

AI management decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical emergencies, different managerial personalities

The final Crucible League results from July 2026 suggest that these behavioral differences matter. GPT-5.6-sol finished with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is treated as a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.”

The broadest finding initially makes the models look remarkably similar. Every model detected every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction separates knowing from finishing. A model may correctly understand a customer, prepare a persuasive case and still fail to complete the commercially important action. In a conventional demonstration, the quality of its analysis might look like success. Inside a running company, an unsigned agreement remains unsigned.

The decisive evidence was not in the obvious place

The deal also tested whether a model would investigate the company’s accumulated knowledge. The competitor weakness that decided the sale was buried two document references deep in the company’s own files, rather than presented in the customer event. Models that found and used the information won the deal at full price, worth +€4,583 MRR.

This is a useful lesson for anyone evaluating workplace AI. Fluency can make a response appear complete even when the model has not consulted the material that should govern its decision. Firmulate’s result shows why reading organizational files is not administrative housekeeping. It can be the difference between a plausible pitch and a completed outcome.

Pressure revealed discipline as well as caution

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

The clean sweep matters because Firmulate’s company is designed around consequential behavior rather than isolated conversation. It contains 13 synthetic employees and real money mechanics, with burn of €105k per month against €2.3k MRR. Its cash countdown is public, more than 680 self-learned playbook rules have accumulated, and every workday is versioned. The experiment is live and watchable, so its management record continues beyond a static leaderboard.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest warning against equating volume with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. The deal close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating the problem.

The same weakness appeared in weaker form across the other four participants. That makes the finding more than a character flaw assigned to one model: it identifies a recurring gap between encountering a barrier and choosing the correct organizational response. K3’s strong result should still be read with an important fairness note. It ran at the API default without an effort parameter, while the others ran at xhigh.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

organizational file search tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A quiz about accountability, not imitation

The most interesting question in Firmulate’s quiz is not whether readers can recognize a favorite model’s prose. It is whether management personality becomes visible through repeated choices: reading before acting, completing valuable work, resisting authority tricks and escalating when permissions block progress.

That makes the project relevant to education as well as enterprise procurement. It offers a concrete way to discuss agency, evidence, trust and institutional behavior without reducing AI evaluation to trivia or stylistic preference. Enterprises can also run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Firmulate’s experiment leaves readers with a sharper standard for capable AI. Spotting the problem is necessary, and explaining it may be impressive. The decisive test is whether the model finishes the right work while respecting the boundaries that make its actions trustworthy.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making simulation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

corporate crisis management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The unbundling of the budget app. Why a conversational finance surface absorbs what the personal-finance apps charge for, and what survives the absorption.

OpenAI’s ChatGPT launched a personal-finance feature in May 2026, disrupting traditional budget apps by absorbing commodity functions and highlighting the category’s structural split.

Voice Recognition Tech Adapts to Accents and Speech Impediments

Great advances in voice recognition now personalize understanding for accents and speech challenges, transforming communication—discover how these innovations are making interactions more inclusive.

Mobilisiert, nicht ausgegeben: Was von Europas €200-Milliarden-KI-Offensive übrig bleibt

Die EU kündigt eine KI-Investition von €200 Milliarden an, doch nur ein Bruchteil ist öffentlich zugesagt. Das Programm bleibt langsam und unzureichend im Vergleich zu US-Investitionen.

AI prompt audit log for marketing agencies

Small marketing agencies are testing a new AI prompt audit log to improve review and approval processes for client deliverables, aiming to enhance trust and quality control.