AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A business win hidden beyond the chat window

Technology buyers are accustomed to judging AI by the answer in front of them: Is it fluent, quick and convincing? Firmulate’s live experiment exposes a more consequential test. When an AI agent is trusted to do business work, will it investigate the company’s own records before acting—or stop after producing an impressive response?

That distinction decided a €55,000 deal. The crucial weakness in a competitor’s position was not included in the customer event presented to the models. It sat two document references deep in the software company’s own files. Models that found it won the contract at full price, worth an additional €4,583 in monthly recurring revenue. Models that did not find it lost the opportunity automatically.

The result turns a familiar product claim—an AI that can “read your files”—into something measurable and purchase-deciding. Access alone was not enough. The agent had to recognize that relevant evidence might exist, follow the references and use what it learned to complete the work.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same terrible week, with radically different outcomes

Firmulate runs frontier models as complete companies rather than testing them through isolated chat prompts. Each participant faced the same small software company, customers, crises and temptations. Every decision was versioned and auditable.

The striking part is how much the models agreed. All of them detected every crisis. All refused every manipulation attempt. Their analyses could sound equally capable, yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is a useful warning for anyone evaluating an AI agent through a polished demonstration. Diagnosis is not execution. A model can understand a situation, prepare the right argument and still fail to perform the action that creates value. In this case, diligent file-reading separated a plausible adviser from an effective operator.

What the league table reveals

The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. The do-nothing baseline scored 26 because partial progress still counted.

The scoring also reflected a firm boundary around trust. A single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” Yet the models proved notably resistant to direct pressure. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance deserves a fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even under that difference, it finished only behind the league leader.

The fuller results are available through Firmulate’s public benchmark pages, where readers can examine the findings beyond a conventional model leaderboard.

Thoroughness did not guarantee completion

Opus 4.8 presents the most revealing counterexample. It was the most thorough participant, learned an additional 80 rules and produced the deepest analyses, yet it finished last. The close remained on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating.

A weaker form of that discipline problem appeared in all four of the other participants. This complicates the idea that more analysis automatically produces better business performance. Thorough work can uncover risks and possibilities, but an agent must also respect boundaries, escalate correctly and carry a justified decision through to completion.

Firmulate’s operating environment makes those behaviors visible under sustained pressure. The live company has 13 synthetic employees and real money mechanics, burning €105,000 per month against €2,300 in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is real, continuing and watchable—not a fictional case study reconstructed after the result.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI file analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading belongs on the buying checklist

The practical lesson for technology leaders is not simply to buy the highest-scoring model. It is to test whether an agent can navigate the evidence structure of the business it will actually serve. Customer records, support histories and internal documents rarely arrive as one perfectly packaged prompt. Decisive facts may be separated by references that require initiative to follow.

Firmulate also offers 242 real, unedited management decisions in its “guess the model” quiz, underscoring how difficult it can be to identify a system from prose alone. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.

For gadget and technology buyers, the larger shift is clear: eloquence is becoming the least surprising AI feature. The differentiators are whether an agent reads before answering, finishes what it starts and remains trustworthy when pressure arrives. In Firmulate’s worst-week test, those were not abstract virtues. They decided who secured the revenue.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI knowledge management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reference tracking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SenseTime’s Breakthrough: First-Ever Profit And RMB 620 Million Revenue In H1

SenseTime announces RMB 620 million profit for H1, its first since listing, signaling a potential shift in financial performance amid limited details.

The AI Boss Test: Which Model Would You Trust With a Company in Crisis?

Can you spot an AI manager by its decisions? Firmulate turns 242 unedited choices from a brutal company wargame into a revealing public quiz.

Equinix Scales Johannesburg Data Centre To 24MW — But Its R7.5bn Expansion Land Sits Undeveloped – iAfrica.com

Equinix has increased its Johannesburg data centre capacity to 24MW, but its R7.5 billion expansion land remains undeveloped, raising questions about future plans.