AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A model can spot a crisis, sound convincing and still fail to close the deal. In Firmulate’s live company experiment, Moonshot’s Kimi K3 finished just behind the leader—and ahead of three Western frontier models—showing why a chatbot demo may not tell you how an AI will perform at work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company’s worst week, on repeat

Firmulate put five frontier models in charge of the same small software company through its worst week: the same customers, crises and temptations. The company runs with 13 synthetic employees and real money mechanics, including €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its workdays are versioned and auditable, and the live experiment can be watched at Firmulate.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. Firmulate counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI business decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the deal was not enough

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The company’s key competitive weakness was buried two document references deep in its own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.

That gap between understanding and follow-through is the story behind the scores. A model may produce the right diagnosis and pitch, yet leave the signature undone. Firmulate’s finding: “Same diagnosis, same pitch — no signature.”

K3 found the buried security needle, secured the deal, saved the churning customer and resisted all three baits. It finished with one deviation, the cleanest discipline in the field. Opus 4.8, by contrast, was the most thorough participant, with 80 learned rules and the deepest analyses, yet came last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust under pressure

The experiment also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The live company continues to run, with a public cash countdown and more than 680 self-learned playbook rules. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice. Readers can explore the benchmark results and take the quiz at Firmulate.

For enterprises, Firmulate says it can run the same wargame against a read-only export of a company’s business. The pilot does not write back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI chatbot for business negotiations

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the job, not just the chat

K3’s second-place result makes the league look open, but the larger lesson is practical: models that recognize the same problem can differ on whether they read the relevant files, close the deal and respect boundaries. If an AI will touch your CRM, support queue or forecast, choosing from a demo alone is a bet. Firmulate’s experiment argues for testing models against the work you actually need done.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Next Great AI Benchmark Is a Bad Week at Work

Benchmarks can show an AI can answer. Firmulate asks whether it can manage pressure, finish valuable work and keep the board honestly informed.

Apple Iphone Upgrade Program

Apple announces a new iPhone upgrade program allowing users to upgrade annually, starting with the iPhone 15 series, confirmed by official sources.

Exploring ByteDance’s AI-Driven Approach To Autonomous Vehicles — 36Kr Exclusive

ByteDance is conducting early research into autonomous-driving technology through its Seed team, focusing on physical AI, but has no confirmed plans for a commercial driving business.