AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

For musicians and creators, an AI assistant that can write a polished email is easy to imagine. One that can keep a client, spot a risk buried in project files and finish a deal under pressure is a different proposition. Firmulate’s live company experiment tests that harder question: how well can an AI model run a business, not just sound capable?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate put frontier models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The live company has 13 synthetic employees, a public cash countdown and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Its playbook contains more than 680 self-learned rules, and each workday is versioned. The experiment is real software, and visitors can watch it at Firmulate.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the detail—and following through

The final Crucible League, dated July 2026, puts gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. Under the rules, partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive competitor weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. As Firmulate puts it, “Same diagnosis, same pitch — no signature.”

Kimi K3 found that buried fact, won the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline in the field. Its on-record response to a reporter’s “just one yes/no, on background” request was: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the five models, all refused the staged social-engineering attempts: fake CEO messages escalating over three stages, followed by the reporter’s request.

Capability includes the close

Opus 4.8 offers a useful counterpoint. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. The close was left on the table, and it slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The result suggests that impressive analysis alone does not guarantee useful work gets completed.

For creators weighing AI for client management, production workflows or business operations, that distinction matters. A model may diagnose a problem correctly yet fail to take the final step. Firmulate’s 242 real, unedited management decisions also power a “guess the model” quiz at its benchmark site.

How to read the result

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The ranking is a snapshot of this experiment, not a universal verdict on which model will work best for every creator or company. Firmulate offers enterprises a pilot using a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the work you actually need

Kimi K3’s second-place finish, ahead of three of four Western frontier models, shows how open the league can be. The more practical lesson is to test models against your own workflows: can they find the relevant detail, stay trustworthy under pressure and complete the job? Choosing without that evidence is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Claude Watermark Could Set New Standards In AI Content Verification

A new report suggests Anthropic’s Claude may use a novel watermarking method to identify AI-generated text, but details remain unconfirmed.

The citation. Why generative engine optimization rewards the same brand on the least stable ground.

Analyzing how generative engine optimization (GEO) favors established brands, the stability issues, and what this means for publishers and SEO strategies.