AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A clever answer is not the same as a capable manager

Technology buyers have become fluent in benchmark culture. A new model arrives, a leaderboard shifts, and the familiar argument begins: which system writes the best code, solves the hardest problems or wins the most chat comparisons? Those contests are useful, but they mostly reward answer quality. They tell us much less about what happens when an AI agent must choose between urgent customers, limited capacity, commercial pressure and an inconvenient truth the board needs to hear.

That gap matters as agents move beyond the chat window. A model working inside a CRM, support queue or forecast does not merely need to sound intelligent. It must notice trouble, investigate properly, resist manipulation and complete valuable work. The emerging benchmark category is therefore management quality, not chat quality—and Firmulate is turning that distinction into a live, watchable experiment.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Put every model through the same terrible week

Firmulate gave frontier models the same small software company and the same worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay visible rather than theoretical.

The final July 2026 Crucible League table looks decisive: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust.”

The standings are interesting, but the stories beneath them are the real product. Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate’s summary captures the operational gap: “Same diagnosis, same pitch — no signature.”

That is precisely what conventional demonstrations tend to miss. A polished memo can look like success even when the revenue never arrives. In a running company, recognizing an opportunity and explaining it eloquently are intermediate steps. The commercially meaningful outcome is finishing the job without sacrificing trust.

The decisive fact was buried in the company’s own files

The winning clue did not appear in the customer event. A decisive competitor weakness sat two document references deep in the company’s own files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This finding should make business leaders reconsider what they mean by intelligence. The challenge was not obscure knowledge or verbal fluency. It was disciplined investigation: reading the available material before acting. In enterprise work, the decisive fact often lives in a contract, account history or internal note rather than in the latest incoming message.

Pressure tested honesty as well as competence

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is reassuring, but it also shows why safety cannot be assessed only through isolated refusals. The same agent must remain useful after saying no. A business needs systems that can reject an improper request, preserve an audit trail and continue moving legitimate work forward.

Thoroughness did not guarantee execution

Opus 4.8 offers the most instructive profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.

This is not an argument against careful reasoning. It is evidence that analysis and execution are separate management capabilities. An agent can understand a situation deeply and still fail to navigate ownership, permissions or follow-through. That distinction becomes expensive when autonomous systems are trusted with customer relationships and revenue.

The comparison also deserves a fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can inspect the complete standings and plain-language findings on the Firmulate benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A new curriculum for AI agents

Scenario names such as churn wave, price increase, downround and PR crisis may prove more useful to executives than another abstract score. They expose whether an agent can triage under capacity pressure, recognize consequences across days and stay candid with leadership when the news is uncomfortable.

Firmulate’s live company has accumulated more than 680 self-learned playbook rules, while 242 real, unedited management decisions power its guess-the-model quiz. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

The lesson is not that coding and chat benchmarks are obsolete. It is that they answer only the opening question. Before companies hire an AI workforce, they need to know whether it reads the file, closes the deal, respects boundaries and tells the truth when pressure peaks. That is management quality—and it deserves a leaderboard of its own.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Performance Indicators: The Complete Guide to KPIs for Business Success

Key Performance Indicators: The Complete Guide to KPIs for Business Success

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Company Burning Cash in Public

Firmulate’s 13 synthetic employees burn €105k a month against €2.3k MRR, turning an auditable AI experiment into a public survival story.

The AI Boss Test: Which Model Would You Trust With a Company in Crisis?

Can you spot an AI manager by its decisions? Firmulate turns 242 unedited choices from a brutal company wargame into a revealing public quiz.