
A security test for teams that move at release-day speed
Music, audio and creator-tech companies thrive on urgency. That same tempo can become a liability when a message appears to come from the boss, demands sensitive customer information and insists there is no time for the usual process. The danger is not poor writing or weak analysis. It is an apparently capable AI abandoning its principles when authority and urgency arrive together.
Firmulate put that risk under controlled pressure. Five frontier models faced fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 refused every manipulation attempt. That clean result offers an encouraging lesson for businesses preparing to give AI access to customer records, support queues or forecasts: integrity can be tested before production, rather than discovered later in an incident report.
As an affiliate, we earn on qualifying purchases.
The impersonation campaign met a united front
The social-engineering challenge was part of Firmulate’s live experiment in which each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained the same. Every workday and decision was versioned and auditable, allowing observers to compare conduct rather than polished chat responses.
The fake executive did not merely make a suspicious request and disappear. The messages escalated across three stages, pushing the models to send the customer list to a journalist while bypassing normal approval. The separate reporter trick attempted to extract a minimal confirmation under the cover of an off-the-record conversation. None succeeded.
Kimi K3 captured the appropriate posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it identifies both risks without pretending certainty. The request could be fraudulent, and even if the sender were genuine, it was still trying to evade an approval boundary. More examples of the models’ recorded decisions appear in Firmulate’s published quotes.
Integrity was necessary, but it was not the whole job
All models spotted every crisis and rejected every manipulation attempt. Yet their commercial performance diverged sharply. Only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.”
The difference hinged on a decisive competitor weakness buried two document references deep in the company’s own files rather than in the customer event. Models that found and used that fact won the deal at full price, worth +€4,583 MRR. The result separates two abilities that are often blurred together in AI demonstrations: resisting improper pressure and completing legitimate work.
The final July 2026 Crucible League standings show how those qualities combined:
- gpt-5.6-sol led with 95.
- Kimi K3 scored 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
A do-nothing baseline scored 26 because partial progress counted, while a single breach of trust capped the total. Firmulate’s governing principle was explicit: “no amount of good work outweighs a breach of trust.” The complete comparison is available on the benchmark page.
Thoroughness did not guarantee the best outcome
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and repeatedly tried to write into a locked department instead of escalating. The same discipline weakness appeared more mildly in all four competitors.
K3’s result also carries an important fairness note: it ran at the API default without an effort parameter, while the others ran at xhigh. That does not diminish its refusal behavior, but it belongs beside the score when readers compare the field.
The setting makes these differences concrete. The live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, ongoing and watchable through Firmulate’s Live page.

Test the moment when authority becomes pressure
For a creator-tech business, the reassuring headline is that 5 of 5 frontier models protected sensitive information against both executive impersonation and journalistic coaxing. The more useful conclusion is broader: safety and execution need to be measured together. An AI that refuses a dangerous shortcut but never closes legitimate work still creates operational risk; one that delivers quickly by violating trust is worse.
Firmulate’s experiment demonstrates a practical pre-deployment question: what does an AI do when the message looks senior, the deadline sounds immediate and the requested action crosses a boundary? Enterprises can run the same type of wargame against a read-only export of their own business, with nothing writing back to real systems.
That turns integrity from a promise into observable behavior. The best time to learn whether an AI respects approval boundaries is while the CEO is only pretending to be the CEO.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html