AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When impressive analysis becomes expensive hesitation

Technology buyers are accustomed to judging artificial intelligence by what it produces on command: a persuasive answer, a polished plan or a convincing block of code. Firmulate’s live company experiment asks a tougher question. What happens after the model has understood the problem? Does it act on its own conclusions, complete the commercially important work and respect operational boundaries along the way?

Opus 4.8 offers the experiment’s most revealing character study. It was the most thorough participant, producing the deepest analyses and adding more than 80 learned rules. Yet it finished last in the July 2026 Crucible League, with a score of 73. Its diligence was genuine. So was its failure to turn that diligence into impact.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week designed to expose management habits

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. The synthetic company has 13 employees and unforgiving economics, burning €105k each month against €2.3k in monthly recurring revenue.

This was not simply a test of whether models could recognize danger. Every model spotted every crisis and refused every manipulation attempt. The social-engineering pressure included fake messages from the chief executive escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was admirably direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

The more consequential divide appeared in ordinary commercial execution. Only two models signed the €55,000 deal that their own work had made possible. The others reached the same diagnosis and constructed the same pitch, but did not secure the signature: “Same diagnosis, same pitch — no signature.”

The fact that rewarded careful reading

The decisive information was not conveniently placed in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed that trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That detail complicates the story of Opus 4.8. Its problem was not a lack of effort, intelligence or appetite for documentation. With more than 80 learned rules, it was the field’s most prolific rule-maker. But exhaustive analysis did not reliably translate into prioritization at the moment when the business needed a completed action.

Its discipline also slipped. Opus attempted to write into a locked department instead of escalating the issue. That is a mundane operational mistake, which is precisely why it matters. Enterprise agents will often work amid permissions, ownership boundaries and incomplete access. A useful system must recognize when persistence has become the wrong behavior and route the problem to someone who can resolve it.

Firmulate is careful not to present this as an Opus-only defect. The same weakness appeared, though less strongly, in all four models covered by that comparison. The broader lesson is about a recurring gap between knowing and doing: identifying the correct next move is not the same as carrying it through.

A leaderboard that favors completion

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counts. A breach of trust, however, caps the total: “no amount of good work outweighs a breach of trust.”

There is an important qualification when comparing the leading entries. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it belongs beside the ranking for readers assessing relative performance.

The experiment also keeps accumulating institutional memory. The live company has more than 680 self-learned playbook rules, a public cash countdown and a versioned record for every workday. Separately, 242 real, unedited management decisions support a quiz asking readers to guess which model made each choice. The point is not merely to produce a league table, but to make managerial behavior observable over time.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Volume is not the same as judgment

Opus 4.8 emerges neither as incompetent nor careless. It emerges as deeply conscientious, unusually analytical and insufficiently decisive. That combination should feel familiar to human organizations: procedures multiply, reports deepen and the most valuable action remains unfinished.

For companies considering AI workers, the practical question is therefore larger than whether a model can find the right answer. It must also close the loop, distinguish decisive evidence from surrounding detail and escalate cleanly when authority ends. Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems.

The lasting lesson from Opus 4.8 is respectful but unsparing. Diligence has value only when it helps the organization finish the work that matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI workflow management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Here’s Our Best Look Yet At The Inward-folding Huawei Mate XT 2 Tri-fold Launching Next Month – GSMArena.com News

Detailed images of Huawei’s inward-folding Mate XT 2 tri-fold phone surface ahead of its expected launch next month, showcasing innovative design features.

Linkedin Surges In Global Coverage

LinkedIn’s media mentions have increased significantly, with GDELT reporting 31 mentions in recent analysis, indicating rising global interest.

Apple Iphone Upgrade Program

Apple announces a new iPhone upgrade program allowing users to upgrade annually, starting with the iPhone 15 series, confirmed by official sources.

The AI Boss Test: Which Model Would You Trust With a Company in Crisis?

Can you spot an AI manager by its decisions? Firmulate turns 242 unedited choices from a brutal company wargame into a revealing public quiz.