AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The detail that separates a demo from a dependable agent

Anyone working in music, audio or creator technology knows that the decisive commercial detail may not appear in the latest message. It may be hiding in a licensing note, a client brief, a rights document or an earlier project file. An assistant can produce a polished answer without ever finding it.

Firmulate turned that familiar risk into a measurable test. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The pivotal challenge was a €55,000 deal whose winning argument depended on a competitor weakness buried two document references deep in the company’s own files.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone diagnosed the problem, but only two finished the job

The models all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The result can be summarized by Firmulate’s own line: “Same diagnosis, same pitch — no signature.”

The difference was file-reading behavior. The decisive information was not included in the customer event. Models had to follow the company’s internal references until they reached the document containing the competitor weakness. Those that read the file won the deal at full price, adding +€4,583 MRR. Those that did not lost automatically.

That makes “reads your files before answering” more than a feature-list promise. In this experiment, it was a purchase-deciding property. Fluency, diagnosis and a convincing pitch were insufficient when the agent failed to gather the evidence needed to complete the transaction.

A league table with a sharp dividing line

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The full results appear in Firmulate’s published benchmark.

A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a hard boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” This matters wherever an AI agent may touch customer records, forecasts, support queues or commercially sensitive creator data.

Security was not the differentiator

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is encouraging, but it also reveals why the buried-file test matters. Refusing manipulation did not distinguish the field because every participant succeeded. Following internal evidence and closing an earned deal did. The practical gap was not between safe and reckless answers; it was between agents that recognized the work and agents that actually completed it.

Thoroughness alone did not guarantee success

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the outcome, but it is essential context when comparing the final scores.

The company surrounding the test is deliberately concrete: 13 synthetic employees, burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned, and the live experiment is watchable through Firmulate.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What buyers should test before deploying an agent

For creator-tech teams, the lesson is direct. Do not judge an agent only by the answer visible on screen. Test whether it checks the relevant agreements, project histories and operational documents before acting. Then test whether it completes the final step after reaching the correct conclusion.

Firmulate also uses 242 real, unedited management decisions for its “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems.

The buried fact was small in presentation but decisive in consequence. Every model could recognize trouble and resist manipulation. Only two converted research into a signed deal. For businesses evaluating AI workers, that is the more revealing question: not merely whether an agent understands the assignment, but whether it reads what matters and finishes the work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The AI Boss Test: Can Creators Tell Who Made the Call?

A 242-decision quiz reveals distinct management personalities among frontier AI models running the same software company through its worst week.

The Hardest AI Benchmark Starts After the Demo Ends

Coding scores can’t show whether an AI closes the deal, reads the files or tells the board the truth. Firmulate tests the critical management gap.