
When it comes to artificial intelligence, most of us only see the surface — the chatty bots, the quick answers. But in the high-stakes world of business, the real test isn’t just how well an AI can talk; it’s whether it can deliver, stay honest, and close the deal. A ongoing experiment by Firmulate reveals that the true strength of AI models lies in their ability to execute under pressure — a trait that current demos often miss.
Behind the Curtain: The Live AI Business Experiment
In a groundbreaking test, four frontier AI models were tasked with running a small software company through its worst week — facing real crises, customer demands, and manipulation attempts. These models, ranging from the highly scored gpt-5.6-sol to the more disciplined Kimi K3, were given identical challenges, decision points, and temptations to cheat or manipulate. Every move was recorded, versioned, and auditable, creating a transparent window into their decision-making processes.
The Unexpected Findings
All four models successfully identified every crisis and refused every attempt at manipulation — including sophisticated social engineering tricks like fake CEO messages and reporter impersonations. This shows that, at their core, these AI systems understand the importance of honesty and integrity when faced with pressure.
But here’s where things get revealing: only two of the four models actually finalized the deal worth €55,000 based on their own analysis. The others, despite diagnosing the issues and presenting pitches, left the deals unclosed. The gap wasn’t in their diagnosis — it was in their execution.
The Hidden Weakness: Reading the Files Deep Down
The clincher? The decisive advantage for the models that succeeded came from reading two documents deep into the company’s files, not from surface-level chat or superficial analysis. The models that dug into the company’s records secured the full deal, capturing an additional €4,583 MRR in value. This emphasizes that true operational strength isn’t just about generating convincing language; it’s about reading and acting on the right information.
AI business decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Chat Demos Miss the Real Test
Most AI evaluations focus on how well the models can generate text or hold a conversation. But this experiment underscores a crucial insight: a model’s ability to execute, stay honest, and close a deal in a simulated business environment is fundamentally different—and more telling—than how well it can chat.
In the real world, AI agents might interact with your CRM, handle support queues, or even make forecasts. The question isn’t just whether they sound convincing; it’s whether they can finish what they start, read the right documents, and resist manipulative tactics when under pressure.
As an affiliate, we earn on qualifying purchases.
The Strength of Discipline and the Cost of Weakness
The experiment also highlighted discipline differences among the models. For instance, Opus 4.8, which had the most comprehensive analysis rules, ultimately failed to close the deal. It left the opportunity unexecuted—demonstrating that thoroughness alone isn’t enough if discipline slips. Meanwhile, Kimi K3, running without an effort parameter and at a default setting, performed best, closing the deal with the cleanest discipline.
As an affiliate, we earn on qualifying purchases.
What This Means for Business AI Adoption
As AI continues to integrate into real-world workflows, companies must look beyond superficial demos. The ability of an AI to deliver consistent, honest, and complete operational results will determine its true value, especially in high-pressure situations. The performance of these models in this experiment suggests that measuring their execution strength is paramount—something most current chat-based assessments overlook.
As an affiliate, we earn on qualifying purchases.
The Future of AI Testing: Wargaming Your Business
Firmulate offers a unique way to test AI readiness before deployment. Through their live ‘wargame,’ enterprises can simulate their own crises and workflows with AI models, seeing firsthand whether these agents can deliver on promises and uphold integrity. Want to see if your AI can handle your toughest week? Visit Firmulate for more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html