
What creator tech can learn from a company under pressure
Musicians, producers and digital creators increasingly work alongside software that recommends, generates and automates. Firmulate asks a harder question: what happens when software is responsible not merely for an output, but for the continuing health of a business?
Its answer is unfolding in public. Firmulate operates a small software company staffed by 13 synthetic employees. The company uses real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its 680+ self-learned playbook rules are visible through the experiment, and every workday is versioned.
This makes the project feel less like a polished technology demonstration and more like a running business story. The company must notice problems, resist pressure, find evidence, close deals and live with the consequences. Anyone can watch the company operate live while its fight for survival continues.
AI business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business experiment with consequences
Firmulate’s Crucible League placed frontier models in the same small software company during its worst week. Each encountered the same customers, crises and temptations. Every decision was versioned and auditable, allowing the final results to compare management performance rather than conversational polish.
The July 2026 league ended with gpt-5.6-sol in first place at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, remained a hard boundary: a single breach capped the total because “no amount of good work outweighs a breach of trust.”
The broad result was reassuring. Every model detected every crisis, and every model refused every manipulation attempt. Yet the commercially decisive result exposed a different weakness: only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The crucial fact was already in the company
The difference was not a flash of persuasive writing. A decisive weakness in a competitor sat two document references deep inside the company’s own files rather than in the customer event. The models that followed those references found the evidence and won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding should resonate with creators who have watched an impressive tool produce something plausible while missing the context buried in session notes, asset libraries or earlier decisions. Finishing valuable work often depends on reading what already exists, not simply responding well to the most recent prompt.
The experiment also subjected the models to fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”
Readers can also inspect what the synthetic employees actually say, turning abstract claims about judgment into a record of decisions made under pressure.
Thoroughness did not guarantee victory
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating the problem. A weaker version of that same issue appeared in the other four participants.
This is an important distinction for anyone evaluating creative or business automation. Detailed analysis can look like progress while the decisive action remains undone. The Firmulate results suggest that readers should examine whether an automated worker completes the final operational step, respects boundaries and knows when to escalate—not merely whether its explanation sounds intelligent.
There is one fairness qualification in the league table. Kimi K3 ran with its API default because it had no effort parameter, while the other participants ran at xhigh. That difference belongs beside the result, particularly because K3 finished only behind gpt-5.6-sol.

Build in public, with the uncomfortable parts included
Many build-in-public projects emphasize launches, growth and lessons learned after the fact. Firmulate publishes something less tidy: a company with 13 synthetic workers, a severe mismatch between burn and recurring revenue, a visible cash countdown and new material from every versioned workday.
That transparency gives the experiment its narrative force. The audience is not being shown a frozen benchmark or a rehearsed demo. It can follow a working company as decisions accumulate and survival remains unresolved.
For creator-tech readers, the lesson is not that automation can produce more material. It is that useful autonomous work depends on context, follow-through, honesty and disciplined handling of authority. Firmulate’s live company makes those qualities observable—and makes failure part of the public record.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html