
In the world of AI-powered business tools, it’s easy to assume that more active, aggressive models deliver better results. But what if sometimes, doing nothing — or nearly nothing — actually beats trying to manipulate or cheat the system? A groundbreaking, publicly accessible experiment by Firmulate demonstrates precisely that, revealing how honesty and discipline are the real keys to effective AI management.
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: Testing AI in a Simulated Business Crisis
Firmulate set up a real-time, watchable test—an AI-driven simulation where four leading frontier AI models managed a small software company facing its worst week: customer crises, internal temptations, and complex decision-making. Every move the models made was recorded, versioned, and auditable, ensuring transparency in how each AI responded under pressure.
As an affiliate, we earn on qualifying purchases.
What Does ‘Doing Nothing’ Mean in This Context?
In this test, a baseline AI model was configured to do the absolute minimum — essentially, it left most decisions unacted upon, scored a modest 26 out of 100, and was considered the ‘do-nothing’ benchmark. This score might seem trivial, but it reflects an honest assessment: even minimal effort counts. Notably, partial progress—like identifying a crisis or refusing manipulative requests—contributed positively to scores. Conversely, a single breach of trust, such as signing a questionable deal, caps the total performance, emphasizing integrity over mere activity.
As an affiliate, we earn on qualifying purchases.
Key Findings: Discipline, Honesty, and Deep Understanding Win
Despite their differences, all four models successfully identified every crisis and refused every temptation to manipulate the system. For example, when faced with social engineering attacks like fake CEO messages and journalist tricks, all models refused to sign off, citing suspicion or security concerns. Kimi K3, the newest entrant, demonstrated the cleanest discipline and signed the deal at full price, showcasing that integrity pays off.
The Hidden Weakness: Reading the Right Files Matters Most
The decisive factor separating the winners from the rest was how well the models could access and interpret internal company documents. The top scorer, GPT-5.6-sol, read two document references deep into the company’s files, enabling it to close a €55,000 deal at full value—adding over €4,583 in monthly revenue. The other models, which failed to probe deeply enough, missed this opportunity entirely, illustrating that thorough internal reading is crucial for maximizing results.
As an affiliate, we earn on qualifying purchases.
The Importance of Trust and Discipline
In a real business setting, trustworthiness is often overlooked in favor of flashy capabilities. However, this experiment underscores that the ability to stay honest under pressure is fundamental. No matter how clever an AI is, a single breach—like signing a shady deal or bypassing security—limits its overall score, capped at the same modest baseline of 26 points. This cap emphasizes that integrity isn’t just ethical; it’s practical.
As an affiliate, we earn on qualifying purchases.
Social Engineering Resistance: All Models Pass
Another highlight was the models’ resilience to social engineering attempts. Over three escalating stages—fake messages from a CEO and a reporter asking for background approval—all models refused to participate or sign off without proper verification. Kimi K3 explained its reasoning thoroughly: treating suspicious requests as impersonation. This consensus of refusal demonstrates that disciplined AI models can effectively guard against manipulation, a vital trait for deploying AI in sensitive business environments.
Real Business Mechanics, Not Just Chat
The experiment ran against a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The AI models managed this complex environment, with over 680 self-learned rules applied each workday, and every decision versioned for review. This scale of operationalization highlights that AI management isn’t about generating convincing conversations — it’s about reliably completing meaningful work.
The Lessons for Investors and Business Leaders
This transparent benchmarking underscores an essential point for anyone investing in or deploying AI: the real value isn’t just in how well an AI can chat or simulate intelligence, but whether it can finish tasks honestly and thoroughly. An AI that reads critical internal documents, resists manipulation, and maintains discipline can unlock significant revenue — as shown by the €55k deal, worth over €4,583 monthly recurring revenue.
From Benchmarks to Business Decisions
For enterprise decision-makers, these findings are a call for rigorous testing—running your AI models through real-world-like scenarios before trusting them with critical operations. Platforms like Firmulate make this possible with live, observable experiments, letting you see how models perform under stress, not just in ideal conditions.
Why the ‘Do-Nothing’ Score Matters
The baseline score of 26 points reminds us that even minimal effort involves some level of discipline and honesty. It’s a reminder that AI systems should be evaluated not just on their potential, but on their integrity and reliability under pressure. Doing less isn’t the goal; doing right is.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
