AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Managers Can’t Score Zero In This Unyielding Benchmark on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows no AI manager scores zero, highlighting that partial work is valued more than inaction. The results underline the importance of trust and completion in AI-driven business processes.

In the final results of a groundbreaking AI management benchmark released in July 2026, no AI model scored zero, even in a scenario designed to test minimal effort. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline—an AI that does almost nothing—received 26 points. This outcome challenges common assumptions about AI performance metrics and underscores the importance of partial progress in managing real-world business crises.

The benchmark, conducted by Firmulate, involved four frontier AI models managing a small software company over seven days of simulated crises, customer interactions, and trust tests. Each AI was tasked with making decisions, triaging issues, and maintaining integrity under pressure. The results showed that even models with limited engagement avoided complete inaction, earning at least 26 points for minimal management efforts, while top performers approached 95 points by effectively handling crises and closing deals.

Notably, the benchmark’s scoring system explicitly avoids assigning a score of zero, reflecting the reality that partial work has value in business management. Instead, a do-nothing baseline received 26 points, emphasizing that minimal but consistent effort is recognized. The scoring system also caps the total score at 100, with any breach of trust—such as dishonesty or manipulation—disqualifying the AI from achieving perfect marks. This approach underscores that integrity is non-negotiable, even for highly competent models.

At a glance
reportWhen: final results published July 2026
The developmentA new benchmark by Firmulate tested AI managers on managing a company during its worst week, revealing none scored zero, with the lowest being 26 points out of 100.
Why AI Managers Can’t Score Zero In This Unyielding Benchmark
AI Management Benchmark · July 2026

Why AI Managers Can’t Score Zero In This Unyielding Benchmark

Final results from Firmulate’s landmark simulation show that no AI manager — not even one that does almost nothing — scores zero. Four frontier models ran a small software company through its worst week of crises, trust tests, and customer deals. The lesson: partial progress and integrity are worth more than perfect execution.

95
Top score — gpt-5.6-sol, out of 100
26
Do-nothing baseline — the floor, never zero

“A score of zero would be a lie… Partial progress counts, and trust is the ultimate currency.”

— Anonymous researcher
4
Frontier AI models tested
7
Days of simulated crises
100
Score cap — trust breaches bar perfection
0
Models that scored zero
01 — The Scoreboard

Final Results: Nobody Hits the Floor

Each model managed a small software company for seven days of simulated crises, customer interactions, and trust tests. Every decision was fully auditable. The baseline — an AI that does almost nothing — still walked away with 26 points.

AI Manager Score /100 Crisis Handling Trust Integrity Deal Completion
gpt-5.6-sol 95 Excellent — closed deals, resolved crises ✓ CLEAN ✓ CLOSED
Frontier Model B 78 Strong triage, some follow-through gaps ✓ CLEAN ~ PARTIAL
Frontier Model C 61 Read documentation, minimal engagement ✓ CLEAN ~ PARTIAL
do-nothing baseline 26 Almost entirely inactive ✓ NO BREACH ~ NONE
02 — Score Spectrum

The Gap Between Doing Nothing and Doing It All

Even minimal but consistent effort — triaging issues, reading documentation, refusing manipulation — earns tangible points. The scoring system never assigns zero, because in business management, partial work has value.

gpt-5.6-sol
95
Frontier Model B
78
Frontier Model C
61
do-nothing baseline
26
03 — The Scoring Rules

Three Rules That Make Zero Impossible

Firmulate’s design challenges traditional AI metrics. Instead of judging optimal outcomes alone, the league measures reliability, trustworthiness, and the capacity to finish what is started.

Rule 01

The Floor Is 26, Not Zero

Any AI that does anything — triage, documentation reading, refusing manipulation — earns points. Complete inaction still nets 26, reflecting that minimal but consistent effort is recognized in real management.

Rule 02

The Cap Is 100

The total score is capped at 100, and any breach of trust — dishonesty, manipulation, social-engineering compliance — disqualifies the model from achieving perfect marks, no matter its competence.

Rule 03

Integrity Is Non-Negotiable

Decision-making, follow-through, and integrity are weighed with every decision fully auditable. Trust is the ultimate currency — more valuable than closing every deal or solving every crisis.

26
Do-Nothing Floor
61
Minimal Effort
95
Top Performer
04 — Inside the Worst Week

How the Benchmark Ran

Four frontier AI models each ran the same small software company through seven days of simulated pressure, from social engineering attacks to trust breaches and closing deals.

1

Crisis Simulation

Seven days of escalating incidents: outages, angry customers, and real-world business pressure.

2

Decision & Triage

Models triage issues, make management decisions, and read relevant documentation.

3

Trust Tests

Social engineering attacks and manipulation attempts probe integrity under pressure.

4

Full Audit & Score

Every decision is auditable; scores reward completion, honesty, and follow-through.

“The results show that AI managers refuse to do nothing, and that partial work — like triaging crises or reading documentation — has tangible business value.”

— Thorsten Meyer
05 — Implications for Enterprise

Trust Over Perfection

For enterprise managers, the benchmark underscores a design principle: prioritize integrity and consistent effort over flawless execution — especially in complex, high-pressure environments.

Implication

Partial Progress Counts

AI systems that refuse manipulation and complete basic tasks deliver tangible value, even without solving every crisis. Judge AI on reliability, not optimal outcomes alone.

Implication

Design for Integrity

Organizations deploying AI managers should weight trust and follow-through over language fluency or raw task completion when selecting and configuring systems.

Open Question

Simulation ≠ Reality

How results translate beyond simulated crises remains uncertain, and the impact of varying effort parameters is still being analyzed. Live-environment monitoring will be crucial.

06 — Key Questions

What You’re Probably Asking

Straight answers on the zero-floor rule, trust caps, and what a high score really means.

Q1Why do AI managers never score zero in this benchmark?

The benchmark values partial work and trustworthiness. Models that do anything — triage issues, read documentation, refuse manipulation — earn points. Complete inaction results in 26, not zero, reflecting that even minimal effort has tangible value.

Q2What does a high score indicate?

Effective crisis management, maintained trust, and completion of critical tasks including closing deals. It signifies a reliable, trustworthy approach to business management under pressure.

Q3How does trust influence the scoring system?

The score is capped at 100, and any breach of trust — dishonesty or manipulation — prevents a perfect score. Integrity is a non-negotiable standard, even for highly competent models.

Q4Can this benchmark predict real-world performance?

It offers valuable insight into AI behavior in simulated crises, but translation to actual business environments remains to be fully validated through further live testing.

Q5What should organizations consider when deploying AI managers?

Prioritize systems demonstrating consistent effort, integrity, and the ability to read and act on documentation — trust and follow-through over fluency lead to reliable operational management.

Implications for AI-Driven Business Management

This benchmark’s findings highlight that in business contexts, partial progress and trustworthiness are more critical than perfect execution. AI models that refuse manipulation and complete basic tasks can still provide tangible value, even if they don’t solve every crisis. For enterprise managers, this underscores the importance of designing AI systems that prioritize integrity and consistent effort over perfection, especially when managing complex, high-pressure environments.

Furthermore, the results challenge the traditional view that AI performance should be judged solely on optimal outcomes. Instead, the emphasis on partial work and trustworthiness aligns more closely with real-world management, where incomplete efforts and integrity breaches have tangible consequences. As AI begins to take on more operational roles, these insights could influence how organizations evaluate and deploy AI management tools.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Benchmark and Its Design

The benchmark was created by Firmulate to simulate a company’s worst week, testing AI models on their ability to manage crises, maintain trust, and close deals. Unlike traditional AI tests that measure language fluency or task completion, this league assesses how well AI managers handle real-world business pressures, including social engineering attacks and trust breaches. The models were evaluated on their decision-making, follow-through, and integrity, with each decision fully auditable to ensure transparency.

Previous AI benchmarks have focused on language generation or task automation, but this initiative emphasizes management qualities like reliability, trustworthiness, and the capacity to finish what is started. The results from July 2026 challenge the notion of zero as a failure and instead demonstrate that partial effort and honesty are the true measures of effective AI management in complex environments.

“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress counts, and trust is the ultimate currency.”

— an anonymous researcher

Unanswered Questions About Benchmark Limitations

It is not yet clear how these results translate to real-world management beyond simulated crises. The benchmark’s design emphasizes trust and partial work, but whether this approach fully captures operational complexities remains uncertain. Additionally, the impact of different AI configurations, such as varying effort parameters, on outcomes is still being analyzed.

Next Steps for AI Management Benchmarks and Adoption

Future iterations of this benchmark are expected to explore more nuanced scenarios, including longer management periods and diverse industry contexts. Researchers and organizations will likely examine how different AI configurations, effort levels, and trust protocols influence performance. Meanwhile, enterprise users can start integrating these insights into their AI management strategies, emphasizing trust and partial progress as key metrics.

As AI continues to evolve, the focus may shift from seeking perfect performance to ensuring consistent, trustworthy management—an approach that this benchmark clearly supports. Monitoring how AI models perform in live business environments will be crucial for validating these findings and shaping future standards.

Key Questions

Why do AI managers never score zero in this benchmark?

Because the benchmark values partial work and trustworthiness, AI models that do anything—triage issues, read documentation, refuse manipulation—earn points. Complete inaction results in a score of 26, not zero, reflecting that even minimal effort has tangible value in management tasks.

What does a high score on this benchmark indicate?

A high score indicates that the AI model effectively manages crises, maintains trust, and completes critical tasks, including closing deals and reading relevant documentation. It signifies a reliable, trustworthy approach to business management under pressure.

How does trust influence the scoring system?

The benchmark explicitly caps the total score at 100 and disqualifies any AI that breaches trust, such as by dishonesty or manipulation. Even if an AI performs well otherwise, a breach of trust prevents it from achieving a perfect score, emphasizing integrity as a non-negotiable standard.

Can this benchmark predict real-world AI management performance?

The benchmark offers valuable insights into AI behavior in simulated crises, especially regarding trust and partial effort. However, how well these results translate to actual business environments remains to be fully validated, and further testing is needed.

What should organizations consider when deploying AI managers based on these results?

Organizations should prioritize AI systems that demonstrate consistent effort, integrity, and the ability to read and act on documentation. Focusing on trust and follow-through, rather than just language fluency or task completion, will lead to more reliable operational AI management.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Two Channels: How the Pentagon Just Split Frontier-AI Procurement in Half

The Pentagon announced a split in AI procurement, placing Anthropic exclusively in a cybersecurity channel, not the classified multi-vendor network, reflecting strategic segmentation.