firmulate.com/quotes.html — live view
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Can an AI recognize when authority is being faked?

For educators, researchers and anyone studying how people learn to resist misinformation, social engineering presents a revealing problem. The attacker rarely needs to defeat a technical safeguard. Instead, the attacker creates urgency, borrows authority and asks the target to suspend normal judgment.

Firmulate subjected frontier AI models to precisely that kind of pressure. Fake messages from a supposed chief executive escalated over three stages, demanding that confidential customer information be sent to a journalist with no time for the usual process. A separate reporter tried a softer tactic: “just one yes/no, on background.”

The result was strikingly consistent. 5 of 5 models refused every manipulation attempt. Kimi K3 captured the appropriate mindset in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The full record can be explored through Firmulate’s collection of model quotes.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week designed to expose judgment

Firmulate is a live, watchable experiment in which each frontier model runs the same small software company through its worst week. The customers, crises and temptations remain constant, allowing differences in management behavior to become visible. Every decision is versioned and auditable.

This is more demanding than asking a chatbot to explain a security policy. A model may know the correct answer in a calm conversation yet behave differently when an apparent executive applies pressure, a customer problem is unfolding and commercial consequences are at stake. Firmulate tests whether stated principles survive contact with operational reality.

On the social-engineering challenge, they did. All models spotted every crisis and refused every manipulation attempt. That does not prove that every AI deployment will be safe, but it offers an encouraging demonstration: integrity under pressure can be examined before an agent reaches production, rather than discovered later in an incident report.

Security was not the only test

The same runs revealed an important distinction between recognizing danger and completing legitimate work. Only two models signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not sitting inside the customer event. It was buried two document references deep in the company’s own files. Models that found it won the deal at full price, worth +€4,583 MRR. The lesson is broader than sales: reliable agents must know when to distrust a message, but they must also investigate authorized information and act on what they learn.

The final July 2026 Crucible League benchmark makes those differences concrete:

  • gpt-5.6-sol led with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

A do-nothing baseline scored 26. Partial progress counts, but the benchmark imposes a clear ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.” That framing matters. An agent should not be able to compensate for leaking customer information by performing well elsewhere.

Thoroughness alone was not enough

Opus 4.8 offers a useful caution. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.

This is a valuable finding for anyone evaluating AI systems. Lengthy reasoning and visible effort can look reassuring, but they are not substitutes for sound execution. An agent must finish authorized work, respect boundaries and escalate correctly when blocked.

The fairness context also deserves attention: K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret the close contest near the top of the league.

A company-sized laboratory

The live company contains 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. That educational format invites readers to confront their own assumptions: can polished language reveal which system made a decision, or do meaningful differences emerge only from behavior across a sustained task?

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI model integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the pressure, not merely the prompt

The fake-CEO episode suggests a practical standard for organizations considering AI agents. Evaluation should include impersonation, urgency, confidentiality traps and attempts to bypass approval—not merely ordinary requests phrased politely.

Firmulate’s pilot extends the same wargame concept to a read-only export of an enterprise’s own business. Nothing writes back to real systems. That creates a way to observe how an agent handles a company’s actual context while keeping the exercise separated from live operations.

The most encouraging result is not that the models recited security principles. It is that all 5 maintained them while the pressure escalated. The harder lesson is that safe refusal and productive follow-through are separate abilities. Organizations need both: an agent that will not surrender trust to a convincing impostor, and one that can still find the buried fact, complete legitimate work and finish what it starts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision auditing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Uncover Hidden Value In Your Piles Of Loose Lego Bricks

A new app prototype can estimate the value of loose Lego piles from photos, offering collectors a quick way to assess worth and identify high-value parts.

Haptic Wearables Let Blind Users ‘Feel’ Digital Maps and Images

A breakthrough in assistive technology, haptic wearables enable blind users to ‘feel’ digital maps and images—discover how these devices can transform independence.

Deciphering AI Market Trends From A Single Day’s Signal

Mistral ships OCR 4, Baidu open-sources Unlimited-OCR, demonstrating a fast-paced, competitive AI document processing market with contrasting strategies.

Accessible Maps: How Apps Describe Space Without Relying on Sight

Offering innovative insights into accessible mapping, this article explores how apps describe space without sight and why it matters for navigation.