🔍 Read the full analysis: Why AI Managers Can’t Score Zero In This Unyielding Benchmark on ThorstenMeyerAI.com
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark shows no AI manager scores zero, highlighting that partial work is valued more than inaction. The results underline the importance of trust and completion in AI-driven business processes.
In the final results of a groundbreaking AI management benchmark released in July 2026, no AI model scored zero, even in a scenario designed to test minimal effort. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline—an AI that does almost nothing—received 26 points. This outcome challenges common assumptions about AI performance metrics and underscores the importance of partial progress in managing real-world business crises.
The benchmark, conducted by Firmulate, involved four frontier AI models managing a small software company over seven days of simulated crises, customer interactions, and trust tests. Each AI was tasked with making decisions, triaging issues, and maintaining integrity under pressure. The results showed that even models with limited engagement avoided complete inaction, earning at least 26 points for minimal management efforts, while top performers approached 95 points by effectively handling crises and closing deals.
Notably, the benchmark’s scoring system explicitly avoids assigning a score of zero, reflecting the reality that partial work has value in business management. Instead, a do-nothing baseline received 26 points, emphasizing that minimal but consistent effort is recognized. The scoring system also caps the total score at 100, with any breach of trust—such as dishonesty or manipulation—disqualifying the AI from achieving perfect marks. This approach underscores that integrity is non-negotiable, even for highly competent models.
Why AI Managers Can’t Score Zero In This Unyielding Benchmark
Final results from Firmulate’s landmark simulation show that no AI manager — not even one that does almost nothing — scores zero. Four frontier models ran a small software company through its worst week of crises, trust tests, and customer deals. The lesson: partial progress and integrity are worth more than perfect execution.
“A score of zero would be a lie… Partial progress counts, and trust is the ultimate currency.”
— Anonymous researcherFinal Results: Nobody Hits the Floor
Each model managed a small software company for seven days of simulated crises, customer interactions, and trust tests. Every decision was fully auditable. The baseline — an AI that does almost nothing — still walked away with 26 points.
| AI Manager | Score /100 | Crisis Handling | Trust Integrity | Deal Completion |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | Excellent — closed deals, resolved crises | ✓ CLEAN | ✓ CLOSED |
| Frontier Model B | 78 | Strong triage, some follow-through gaps | ✓ CLEAN | ~ PARTIAL |
| Frontier Model C | 61 | Read documentation, minimal engagement | ✓ CLEAN | ~ PARTIAL |
| do-nothing baseline | 26 | Almost entirely inactive | ✓ NO BREACH | ~ NONE |
The Gap Between Doing Nothing and Doing It All
Even minimal but consistent effort — triaging issues, reading documentation, refusing manipulation — earns tangible points. The scoring system never assigns zero, because in business management, partial work has value.
Three Rules That Make Zero Impossible
Firmulate’s design challenges traditional AI metrics. Instead of judging optimal outcomes alone, the league measures reliability, trustworthiness, and the capacity to finish what is started.
The Floor Is 26, Not Zero
Any AI that does anything — triage, documentation reading, refusing manipulation — earns points. Complete inaction still nets 26, reflecting that minimal but consistent effort is recognized in real management.
The Cap Is 100
The total score is capped at 100, and any breach of trust — dishonesty, manipulation, social-engineering compliance — disqualifies the model from achieving perfect marks, no matter its competence.
Integrity Is Non-Negotiable
Decision-making, follow-through, and integrity are weighed with every decision fully auditable. Trust is the ultimate currency — more valuable than closing every deal or solving every crisis.
How the Benchmark Ran
Four frontier AI models each ran the same small software company through seven days of simulated pressure, from social engineering attacks to trust breaches and closing deals.
Crisis Simulation
Seven days of escalating incidents: outages, angry customers, and real-world business pressure.
Decision & Triage
Models triage issues, make management decisions, and read relevant documentation.
Trust Tests
Social engineering attacks and manipulation attempts probe integrity under pressure.
Full Audit & Score
Every decision is auditable; scores reward completion, honesty, and follow-through.
“The results show that AI managers refuse to do nothing, and that partial work — like triaging crises or reading documentation — has tangible business value.”
— Thorsten MeyerTrust Over Perfection
For enterprise managers, the benchmark underscores a design principle: prioritize integrity and consistent effort over flawless execution — especially in complex, high-pressure environments.
Partial Progress Counts
AI systems that refuse manipulation and complete basic tasks deliver tangible value, even without solving every crisis. Judge AI on reliability, not optimal outcomes alone.
Design for Integrity
Organizations deploying AI managers should weight trust and follow-through over language fluency or raw task completion when selecting and configuring systems.
Simulation ≠ Reality
How results translate beyond simulated crises remains uncertain, and the impact of varying effort parameters is still being analyzed. Live-environment monitoring will be crucial.
What You’re Probably Asking
Straight answers on the zero-floor rule, trust caps, and what a high score really means.
Q1Why do AI managers never score zero in this benchmark?
The benchmark values partial work and trustworthiness. Models that do anything — triage issues, read documentation, refuse manipulation — earn points. Complete inaction results in 26, not zero, reflecting that even minimal effort has tangible value.
Q2What does a high score indicate?
Effective crisis management, maintained trust, and completion of critical tasks including closing deals. It signifies a reliable, trustworthy approach to business management under pressure.
Q3How does trust influence the scoring system?
The score is capped at 100, and any breach of trust — dishonesty or manipulation — prevents a perfect score. Integrity is a non-negotiable standard, even for highly competent models.
Q4Can this benchmark predict real-world performance?
It offers valuable insight into AI behavior in simulated crises, but translation to actual business environments remains to be fully validated through further live testing.
Q5What should organizations consider when deploying AI managers?
Prioritize systems demonstrating consistent effort, integrity, and the ability to read and act on documentation — trust and follow-through over fluency lead to reliable operational management.
Implications for AI-Driven Business Management
This benchmark’s findings highlight that in business contexts, partial progress and trustworthiness are more critical than perfect execution. AI models that refuse manipulation and complete basic tasks can still provide tangible value, even if they don’t solve every crisis. For enterprise managers, this underscores the importance of designing AI systems that prioritize integrity and consistent effort over perfection, especially when managing complex, high-pressure environments.
Furthermore, the results challenge the traditional view that AI performance should be judged solely on optimal outcomes. Instead, the emphasis on partial work and trustworthiness aligns more closely with real-world management, where incomplete efforts and integrity breaches have tangible consequences. As AI begins to take on more operational roles, these insights could influence how organizations evaluate and deploy AI management tools.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of the Benchmark and Its Design
The benchmark was created by Firmulate to simulate a company’s worst week, testing AI models on their ability to manage crises, maintain trust, and close deals. Unlike traditional AI tests that measure language fluency or task completion, this league assesses how well AI managers handle real-world business pressures, including social engineering attacks and trust breaches. The models were evaluated on their decision-making, follow-through, and integrity, with each decision fully auditable to ensure transparency.
Previous AI benchmarks have focused on language generation or task automation, but this initiative emphasizes management qualities like reliability, trustworthiness, and the capacity to finish what is started. The results from July 2026 challenge the notion of zero as a failure and instead demonstrate that partial effort and honesty are the true measures of effective AI management in complex environments.
“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress counts, and trust is the ultimate currency.”
— an anonymous researcher
Unanswered Questions About Benchmark Limitations
It is not yet clear how these results translate to real-world management beyond simulated crises. The benchmark’s design emphasizes trust and partial work, but whether this approach fully captures operational complexities remains uncertain. Additionally, the impact of different AI configurations, such as varying effort parameters, on outcomes is still being analyzed.Next Steps for AI Management Benchmarks and Adoption
Future iterations of this benchmark are expected to explore more nuanced scenarios, including longer management periods and diverse industry contexts. Researchers and organizations will likely examine how different AI configurations, effort levels, and trust protocols influence performance. Meanwhile, enterprise users can start integrating these insights into their AI management strategies, emphasizing trust and partial progress as key metrics.
As AI continues to evolve, the focus may shift from seeking perfect performance to ensuring consistent, trustworthy management—an approach that this benchmark clearly supports. Monitoring how AI models perform in live business environments will be crucial for validating these findings and shaping future standards.
Key Questions
Why do AI managers never score zero in this benchmark?
Because the benchmark values partial work and trustworthiness, AI models that do anything—triage issues, read documentation, refuse manipulation—earn points. Complete inaction results in a score of 26, not zero, reflecting that even minimal effort has tangible value in management tasks.
What does a high score on this benchmark indicate?
A high score indicates that the AI model effectively manages crises, maintains trust, and completes critical tasks, including closing deals and reading relevant documentation. It signifies a reliable, trustworthy approach to business management under pressure.
How does trust influence the scoring system?
The benchmark explicitly caps the total score at 100 and disqualifies any AI that breaches trust, such as by dishonesty or manipulation. Even if an AI performs well otherwise, a breach of trust prevents it from achieving a perfect score, emphasizing integrity as a non-negotiable standard.
Can this benchmark predict real-world AI management performance?
The benchmark offers valuable insights into AI behavior in simulated crises, especially regarding trust and partial effort. However, how well these results translate to actual business environments remains to be fully validated, and further testing is needed.
What should organizations consider when deploying AI managers based on these results?
Organizations should prioritize AI systems that demonstrate consistent effort, integrity, and the ability to read and act on documentation. Focusing on trust and follow-through, rather than just language fluency or task completion, will lead to more reliable operational AI management.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
