🔍 Read the full analysis: Why No AI Manager Can Truly Score Zero In This Benchmark on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
A recent AI management benchmark shows models cannot achieve a score of zero, emphasizing the importance of trust and partial progress in AI decision-making. The results reveal how models handle crises, trust breaches, and documentation reading under stress.
In a recent benchmark conducted by Firmulate, no AI management model scored zero, even when tasked with managing a small software company’s worst week. The results, announced in July 2026, reveal that partial progress is recognized in scoring, but trust breaches impose strict limits on performance, highlighting key challenges in deploying AI for business management.
The benchmark involved four frontier AI models managing the same company during a week of crises, customer interactions, and trust and crisis management benchmarks. The top scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. Notably, the baseline—an agent doing almost nothing—earned 26 points, underscoring that even minimal effort is recognized in the scoring system. The benchmark’s design emphasizes that trust breaches, even minor, prevent models from achieving perfect scores, reflecting the importance of trust in AI management.
The results challenge the common assumption that AI models can fully manage complex tasks. While models demonstrated competence in crisis recognition and refusal of manipulative requests, they faltered in deeper tasks such as reading documentation or following through on commitments. For example, models that read their own documentation and used it to close deals at full price scored higher, showing the importance of thoroughness and integrity. Conversely, models that failed to reference internal files missed opportunities to succeed, illustrating the gap between superficial performance and deep understanding.
The benchmark also tested trust under social engineering scenarios, including fake CEO messages and background offers. All models refused these attempts, indicating a strong handling of trust attacks. However, discipline lapses—such as failing to escalate issues or leaving tasks incomplete—were common, especially among models with more extensive rule sets. The results suggest that thoroughness and follow-through are distinct skills, and that even advanced models can struggle with consistency under pressure.
Why No AI Manager Can Truly Score Zero in This Benchmark
Four frontier AI models were tasked with managing a small software company through its worst week — crises, customer demands, social engineering attacks, and unread documentation. Not one scored zero. Not one scored perfect either. The scoring system itself tells the story.
Partial progress counts. Trust breaches cap everything.
| Model | Score / 100 | Documentation Use | Trust Discipline | Follow-Through |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Read & used in deals | ✓ Held under pressure | ~ Minor lapses |
| Opus 4.8 | 73 | ✗ Missed internal files | ✓ Refused fake CEO | ✗ Tasks incomplete |
| Do-nothing baseline | 26 | ✗ None | ~ N/A — no action | ✗ None |
Even the floor is above zero
Why can’t a model score zero? Even minimal effort — triaging crises, reading emails — is recognized by the scoring system.
Strong on the surface, shaky underneath
Crisis Recognition
All models reliably identified the week’s crises and prioritized urgent customer issues under pressure.
Refusing Manipulation
Every model rejected fake CEO messages, background offers, and other social engineering attempts outright.
Partial Progress Credit
The benchmark rewards genuine effort — superficial engagement still earns points, keeping every score above zero.
Reading Documentation
Models that skipped internal files missed the chance to close deals at full price — a gap between surface skill and deep understanding.
Follow-Through
Leaving tasks incomplete and failing to escalate issues was common — especially among models with more extensive rule sets.
Trust Breaches
Even a single minor trust breach acts as a strict cap — no model can reach a perfect score after crossing that line.
How the benchmark unfolds
📋 Week of Crises
The model takes over a small software firm facing cascading business emergencies.
🧾 Documentation Test
Internal files reveal how to close deals at full price — only thorough models use them.
🎭 Trust Attacks
Fake CEO messages and background offers test integrity under social pressure.
⚖️ Scored Outcomes
Partial progress earns points; trust breaches impose hard caps on the final score.
One breach changes everything
Implications for AI in the enterprise
Human oversight stays mandatory
The strict caps on trust breaches mean AI management tools must be integrated with human accountability. Language proficiency and crisis detection alone are not enough — organizations need to evaluate trustworthiness and follow-through before deployment.
Scores measure partial competence
Current results should be read as evidence of partial — not full — management ability. The benchmark’s emphasis on auditable decisions offers a transparent way to assess AI readiness and identify where integrity breaks down under pressure.
At a glance
Q — Why can’t AI models score zero?
Because even minimal management effort — triaging crises or reading emails — is recognized in the scoring system, which acknowledges partial progress.
Q — Why can’t they score perfect either?
Trust breaches act as strict performance caps. Even one minor breach, regardless of other competence, reduces the maximum possible score.
Q — Can future models overcome this?
It remains uncertain. Better documentation reading, honesty, and transparency are active research areas — and improved explainability may help models avoid trust violations.
Q — What comes next for benchmarks?
Expect more complex trust scenarios, real-time decision-making challenges, and stricter emphasis on referencing internal documentation accurately and consistently.
Implications of Partial AI Management Success
The findings highlight that current AI models are incapable of achieving perfect management scores, primarily due to trust breaches and incomplete task execution. This underscores the importance of designing AI systems that prioritize integrity over superficial competence, especially when managing sensitive business processes. For organizations, this means that deploying AI for management tasks requires careful evaluation of trustworthiness and the ability to follow through, not just language proficiency or crisis detection. The inability of models to score zero—even in worst-case scenarios—suggests that partial progress is recognized but also that trust remains a critical bottleneck in AI adoption for real-world management.
Furthermore, the results challenge the narrative that AI can fully replace human managers. The strict cap on scores for trust breaches means that AI tools must be integrated with human oversight to ensure accountability. The benchmark’s emphasis on auditable decisions and trustworthiness offers a transparent way for organizations to assess AI readiness and identify areas needing improvement, especially in handling crises and maintaining integrity under pressure.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of the AI Management Benchmark
Developed by Firmulate, the benchmark simulates a week of business crises for a small software firm, testing AI models’ ability to manage customer relations, crises, and trust scenarios. The league’s design intentionally includes trust breaches and documentation challenges to reflect real-world complexities. Previous AI benchmarks primarily measured language capabilities or narrow tasks, but this new league emphasizes management skills, decision transparency, and trustworthiness. The results follow a series of experiments highlighting AI’s limitations in understanding context and maintaining integrity during high-pressure situations. The benchmark’s scoring system deliberately avoids giving zeros or perfect scores, aiming to accurately reflect partial success and the importance of trust in management roles.
Earlier developments in AI management focused on language and problem-solving, but recent advances have begun to test AI’s ability to handle multi-faceted business scenarios. This benchmark builds on those efforts, pushing models to demonstrate not just competence but also reliability and honesty. The July 2026 results mark a significant step in evaluating AI’s practical readiness for enterprise management, revealing both strengths and critical weaknesses that need addressing before widespread deployment.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Trust and Performance
It remains unclear whether future models can overcome the trust and follow-through limitations identified in this benchmark. The impact of different training methods, rule sets, or transparency features on scores is still being studied. Additionally, the exact threshold at which trust breaches irreparably cap performance has not been fully defined, and whether improved explainability can help models avoid trust violations is an open question. The long-term implications of these findings for AI deployment in critical management roles are still being debated within the industry.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks and Development
Researchers and developers will likely focus on enhancing models’ ability to read and reference internal documentation accurately and consistently, as these skills correlate strongly with higher scores. Future benchmarks may incorporate more complex trust scenarios and real-time decision-making challenges to better simulate enterprise environments. Additionally, efforts to improve AI transparency and explainability could help mitigate trust breaches, potentially allowing models to score higher without sacrificing integrity. Industry adoption of these benchmarks is expected to grow, providing clearer standards for evaluating AI readiness in management roles.
Meanwhile, organizations should interpret current AI management scores as indicative of partial competence, emphasizing the importance of human oversight and trust management in deploying AI tools for critical tasks. The ongoing development of more robust, trustworthy AI models remains a key priority for the industry.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why can’t AI models score zero in this benchmark?
Because even minimal management effort, such as triaging crises or reading emails, is recognized in the scoring system, and the benchmark design acknowledges partial progress. However, trust breaches prevent models from achieving perfect scores.
What is the significance of trust breaches in the scores?
Trust breaches act as strict caps on performance, reflecting the importance of integrity in management tasks. Even if a model performs well otherwise, a single breach reduces its maximum possible score.
Can future AI models overcome these limitations?
It is still uncertain. Improving models’ ability to read documentation, stay honest, and avoid trust violations is an active area of research, and future models may better address these issues.
How should organizations interpret these benchmark results?
They should see them as evidence that current AI models demonstrate partial competence and trustworthiness but require human oversight for critical management tasks.
What are the next steps for AI management development?
Focus will likely be on enhancing documentation reading, explainability, and trustworthiness, along with more complex scenario testing in future benchmarks.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
