
A model can spot a crisis, sound convincing and still fail to close the deal. In Firmulate’s live company experiment, Moonshot’s Kimi K3 finished just behind the leader—and ahead of three Western frontier models—showing why a chatbot demo may not tell you how an AI will perform at work.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, on repeat
Firmulate put five frontier models in charge of the same small software company through its worst week: the same customers, crises and temptations. The company runs with 13 synthetic employees and real money mechanics, including €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its workdays are versioned and auditable, and the live experiment can be watched at Firmulate.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. Firmulate counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Finding the deal was not enough
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The company’s key competitive weakness was buried two document references deep in its own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
That gap between understanding and follow-through is the story behind the scores. A model may produce the right diagnosis and pitch, yet leave the signature undone. Firmulate’s finding: “Same diagnosis, same pitch — no signature.”
K3 found the buried security needle, secured the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. Opus 4.8, by contrast, was the most thorough participant, with 80 learned rules and the deepest analyses, yet came last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust under pressure
The experiment also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The live company continues to run, with a public cash countdown and more than 680 self-learned playbook rules. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice. Readers can explore the benchmark results and take the quiz at Firmulate.
For enterprises, Firmulate says it can run the same wargame against a read-only export of a company’s business. The pilot does not write back to real systems.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

AI chatbot for business negotiations
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the job, not just the chat
K3’s second-place result makes the league look open, but the larger lesson is practical: models that recognize the same problem can differ on whether they read the relevant files, close the deal and respect boundaries. If an AI will touch your CRM, support queue or forecast, choosing from a demo alone is a bet. Firmulate’s experiment argues for testing models against the work you actually need done.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
