
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
An AI can find the money and still leave it on the table
For anyone weighing where automation might touch a company’s revenue, the key question isn’t just whether an AI can spot an opportunity. It’s whether it can carry that opportunity through to a decision. In Firmulate’s business experiment, every model identified a valuable deal. Only two signed it.
The deal was worth €55,000, with a reported value of +€4,583 in monthly recurring revenue. The gap between recognizing it and closing it offers a practical way to think about what AI might do inside a business—and what it might miss.
A company under pressure
Firmulate put frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The experiment is part of a live company simulation that readers can watch at Firmulate.
The simulated company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. A public cash countdown makes the pressure visible. Its playbooks have accumulated more than 680 self-learned rules, and each workday is versioned.
The difference between knowing and doing
In the final Crucible League, published in July 2026, gpt-5.6-sol led with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s stated standard is blunt: “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” A model can identify a promising move and still fail to complete it.
The decisive clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price. That detail makes the outcome less about a polished pitch and more about whether an AI can bring scattered business context to bear on a decision.
Trust under pressure
The models also faced staged attempts to exploit authority and access: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. The deal went unsigned, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a qualification when comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The result should be read with that difference in mind.
From watching to a company-specific pilot
For business leaders, a benchmark can show how models behave in a shared scenario. A company-specific wargame can ask a more direct question: how would an AI respond to a crisis using your business context and playbooks? Firmulate says enterprises can run a pilot from a read-only export, testing scenarios against their own company and receiving a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.
Readers can also explore a “guess the model” quiz built from 242 real, unedited management decisions at Firmulate’s quiz page. The decisions offer a closer look at how the models handled the experiment, beyond their final scores.

Put your own playbooks to the test
The experiment suggests that recognizing a crisis, protecting trust and completing a commercially sound decision are separate tests. A pilot can put those tests against your own company’s scenarios using a read-only export, while keeping real systems untouched. Learn more at Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
