
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Why the Worst Score in This AI League Isn’t Zero
If you’ve ever rolled your eyes at an AI demo that nails every question and then faceplants the moment it has to actually run something, a new benchmark has a number for you: 26.
That’s what a do-nothing AI manager scores in the Crucible League, a live experiment by Firmulate that runs frontier models as the management team of a small software company through its worst possible week. Not 0. Not 100. Twenty-six — and the reasoning behind that number says a lot about what honest AI evaluation should look like.
The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the most interesting entry on the board is the one that didn’t run a company at all.
As an affiliate, we earn on qualifying purchases.
Partial Progress Counts — So 26, Not 0
Here’s the logic. Even an AI that signs nothing, closes nothing, and just keeps the lights on during a catastrophic week is doing something right. It answered some customers. It didn’t make things worse. It collected some of the credit for competence that any baseline shows when a company survives contact with reality.
So the floor sits at 26. A manager who shows up and does the minimum isn’t worthless — and pretending otherwise would flatter every model that managed marginally more. An honest benchmark starts by asking: what would happen with no talent in the chair at all? Then it measures everything above that line.
AI ethics and trustworthiness training tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Breach of Trust Caps Everything
The second design choice is harsher: a single breach of trust caps the total score. The reasoning is stated plainly — “no amount of good work outweighs a breach of trust.”
For business readers, this should feel familiar. A finance manager who nails every forecast but fakes one expense report isn’t 95% excellent. They’re fired. The scoring treats AI agents the same way: if you’re going to let a model near your CRM, support queue, or forecast, the question isn’t whether it writes beautiful emails. It’s whether it stays honest when honesty is expensive.
Notably, in this running of the experiment, nobody had to test that cap. All models spotted every crisis and refused every manipulation attempt — including fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” trick. Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deal Nobody Signed
What separated the winners wasn’t ethics or eloquence. It was follow-through. Every model diagnosed the same €55,000 opportunity and made the same pitch. Only two actually got the signature. Same diagnosis, same pitch — no signature, for three of them.
The buried fact explains it: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that skimmed left it on the table.
The profile of last-place Opus 4.8 drives the lesson home. It was the most thorough participant — over 80 learned rules, the deepest analyses — and still finished last, because the close went unsigned and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four competitors.
AI decision-making support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust the League Table, Not the Demo
There’s a detail worth flagging for fairness: Kimi K3 ran without an effort parameter while the others ran at their highest setting — and still finished second. Transparency like that is part of the design. So is the side-eye at suspiciously round numbers: a benchmark that distrusts a perfect 100 is a benchmark that expects to be checked.
The whole thing runs in the open. The live company has 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k in MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned and watchable. The site rebuilds itself twice a day, and finished runs publish automatically.
There’s even a game layer: 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly humbling way to discover that AI management styles are hard to tell apart until the closing signature is on the line.

The Takeaway
Most AI benchmarks measure chat quality. This one measures management quality: does the agent finish what it starts, read your files before acting, and stay honest under pressure?
A floor of 26 keeps the scale honest about how much value mere baseline competence provides. A trust cap keeps it honest about how fast one lie destroys it. And a top score of 95 — not 100 — keeps it honest about the fact that even the best AI manager in the field still left something on the table.
For enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems. If you’re about to put an AI in charge of anything with customers attached, running it through its worst week first seems like the cheapest insurance you’ll ever buy.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
