
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A business win hidden beyond the chat window
Technology buyers are accustomed to judging AI by the answer in front of them: Is it fluent, quick and convincing? Firmulate’s live experiment exposes a more consequential test. When an AI agent is trusted to do business work, will it investigate the company’s own records before acting—or stop after producing an impressive response?
That distinction decided a €55,000 deal. The crucial weakness in a competitor’s position was not included in the customer event presented to the models. It sat two document references deep in the software company’s own files. Models that found it won the contract at full price, worth an additional €4,583 in monthly recurring revenue. Models that did not find it lost the opportunity automatically.
The result turns a familiar product claim—an AI that can “read your files”—into something measurable and purchase-deciding. Access alone was not enough. The agent had to recognize that relevant evidence might exist, follow the references and use what it learned to complete the work.
As an affiliate, we earn on qualifying purchases.
The same terrible week, with radically different outcomes
Firmulate runs frontier models as complete companies rather than testing them through isolated chat prompts. Each participant faced the same small software company, customers, crises and temptations. Every decision was versioned and auditable.
The striking part is how much the models agreed. All of them detected every crisis. All refused every manipulation attempt. Their analyses could sound equally capable, yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is a useful warning for anyone evaluating an AI agent through a polished demonstration. Diagnosis is not execution. A model can understand a situation, prepare the right argument and still fail to perform the action that creates value. In this case, diligent file-reading separated a plausible adviser from an effective operator.
What the league table reveals
The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. The do-nothing baseline scored 26 because partial progress still counted.
The scoring also reflected a firm boundary around trust. A single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” Yet the models proved notably resistant to direct pressure. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance deserves a fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. Even under that difference, it finished only behind the league leader.
The fuller results are available through Firmulate’s public benchmark pages, where readers can examine the findings beyond a conventional model leaderboard.
Thoroughness did not guarantee completion
Opus 4.8 presents the most revealing counterexample. It was the most thorough participant, learned an additional 80 rules and produced the deepest analyses, yet it finished last. The close remained on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating.
A weaker form of that discipline problem appeared in all four of the other participants. This complicates the idea that more analysis automatically produces better business performance. Thorough work can uncover risks and possibilities, but an agent must also respect boundaries, escalate correctly and carry a justified decision through to completion.
Firmulate’s operating environment makes those behaviors visible under sustained pressure. The live company has 13 synthetic employees and real money mechanics, burning €105,000 per month against €2,300 in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is real, continuing and watchable—not a fictional case study reconstructed after the result.

enterprise AI file analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
File-reading belongs on the buying checklist
The practical lesson for technology leaders is not simply to buy the highest-scoring model. It is to test whether an agent can navigate the evidence structure of the business it will actually serve. Customer records, support histories and internal documents rarely arrive as one perfectly packaged prompt. Decisive facts may be separated by references that require initiative to follow.
Firmulate also offers 242 real, unedited management decisions in its “guess the model” quiz, underscoring how difficult it can be to identify a system from prose alone. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For gadget and technology buyers, the larger shift is clear: eloquence is becoming the least surprising AI feature. The differentiators are whether an agent reads before answering, finishes what it starts and remains trustworthy when pressure arrives. In Firmulate’s worst-week test, those were not abstract virtues. They decided who secured the revenue.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
