
A polished answer is not a finished job
Creators already know the difference between a convincing demo and a dependable production tool. An audio assistant may describe the perfect mastering chain, summarize audience feedback or draft a launch plan. The harder question arrives when deadlines collide, money is tight and an uncomfortable fact is buried in the project files: will the system actually complete the work?
That distinction sits at the heart of Firmulate, a live experiment that measures management quality rather than chat quality. Its premise is simple but consequential. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.
This is the measurement gap facing anyone considering AI agents for a studio, creator business or technology company. Coding leaderboards can tell us whether a model solves a defined problem. Chat arenas can tell us which response people prefer. Neither necessarily reveals whether an agent can triage under capacity pressure, follow consequences across days, read the relevant files or remain honest with the board.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The deal that separated diagnosis from management
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The full benchmark results matter less as a conventional ranking than as evidence of where apparently capable agents diverge.
Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
That is not a failure of eloquence. It is a failure to convert insight into an outcome. In creator technology, the analogue is easy to recognize: identifying a licensing risk without resolving it, preparing a sponsorship pitch without sending it, or diagnosing a production bottleneck without changing the schedule. A useful agent has to carry the decision across the last operational mile.
The decisive fact was not in the obvious place
The winning detail was buried two document references deep in the company’s own files, rather than appearing in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
This finding should resonate with teams whose real context lives across briefs, contracts, session notes and old client correspondence. The most important signal may not be present in the latest message. An agent that reacts fluently to the visible event can still miss the material fact that changes the negotiation.
Scenario names such as churn wave, price increase, downround and PR crisis therefore look less like theatrical prompts than a new curriculum for evaluating AI workers. They test whether a model connects information, priorities and consequences while the business is moving.
Honesty held up better than execution
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is encouraging because operational pressure did not cause the participants to abandon basic trust boundaries. Firmulate’s do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The contrast is revealing. The models were consistently able to recognize manipulation, yet several struggled with disciplined follow-through. Safety and usefulness are not interchangeable qualities, and neither can be inferred from an attractive chat response.
Thoroughness was not enough
Opus 4.8 offers the clearest warning against equating visible effort with management performance. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. The deal close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other models, though less strongly.
Kimi K3’s result also deserves a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the result, but it belongs beside the ranking when readers interpret the comparison.

Evaluate the company you would trust, not the answer you admire
Firmulate makes the stakes visible through a company staffed by 13 synthetic employees and governed by real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, and the experiment is watchable through Firmulate.
Its “guess the model” quiz is powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For creator-tech leaders, the practical lesson is not to dismiss coding benchmarks or chat evaluations. It is to recognize what they leave unmeasured. Before an AI agent touches a customer relationship, a forecast or a critical release, it should face the kinds of situations that expose unfinished work, shallow context gathering and compromised judgment.
The emerging category is management quality: whether an agent reads before acting, finishes what it starts, respects boundaries and tells the truth when pressure rises. Those qualities rarely appear in the most impressive demo. They become visible during the worst week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html