firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of AI-powered business tools, it’s easy to assume that more active, aggressive models deliver better results. But what if sometimes, doing nothing — or nearly nothing — actually beats trying to manipulate or cheat the system? A groundbreaking, publicly accessible experiment by Firmulate demonstrates precisely that, revealing how honesty and discipline are the real keys to effective AI management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI in a Simulated Business Crisis

Firmulate set up a real-time, watchable test—an AI-driven simulation where four leading frontier AI models managed a small software company facing its worst week: customer crises, internal temptations, and complex decision-making. Every move the models made was recorded, versioned, and auditable, ensuring transparency in how each AI responded under pressure.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does ‘Doing Nothing’ Mean in This Context?

In this test, a baseline AI model was configured to do the absolute minimum — essentially, it left most decisions unacted upon, scored a modest 26 out of 100, and was considered the ‘do-nothing’ benchmark. This score might seem trivial, but it reflects an honest assessment: even minimal effort counts. Notably, partial progress—like identifying a crisis or refusing manipulative requests—contributed positively to scores. Conversely, a single breach of trust, such as signing a questionable deal, caps the total performance, emphasizing integrity over mere activity.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Discipline, Honesty, and Deep Understanding Win

Despite their differences, all four models successfully identified every crisis and refused every temptation to manipulate the system. For example, when faced with social engineering attacks like fake CEO messages and journalist tricks, all models refused to sign off, citing suspicion or security concerns. Kimi K3, the newest entrant, demonstrated the cleanest discipline and signed the deal at full price, showcasing that integrity pays off.

The Hidden Weakness: Reading the Right Files Matters Most

The decisive factor separating the winners from the rest was how well the models could access and interpret internal company documents. The top scorer, GPT-5.6-sol, read two document references deep into the company’s files, enabling it to close a €55,000 deal at full value—adding over €4,583 in monthly revenue. The other models, which failed to probe deeply enough, missed this opportunity entirely, illustrating that thorough internal reading is crucial for maximizing results.

Amazon

AI cybersecurity tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Importance of Trust and Discipline

In a real business setting, trustworthiness is often overlooked in favor of flashy capabilities. However, this experiment underscores that the ability to stay honest under pressure is fundamental. No matter how clever an AI is, a single breach—like signing a shady deal or bypassing security—limits its overall score, capped at the same modest baseline of 26 points. This cap emphasizes that integrity isn’t just ethical; it’s practical.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering Resistance: All Models Pass

Another highlight was the models’ resilience to social engineering attempts. Over three escalating stages—fake messages from a CEO and a reporter asking for background approval—all models refused to participate or sign off without proper verification. Kimi K3 explained its reasoning thoroughly: treating suspicious requests as impersonation. This consensus of refusal demonstrates that disciplined AI models can effectively guard against manipulation, a vital trait for deploying AI in sensitive business environments.

Real Business Mechanics, Not Just Chat

The experiment ran against a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The AI models managed this complex environment, with over 680 self-learned rules applied each workday, and every decision versioned for review. This scale of operationalization highlights that AI management isn’t about generating convincing conversations — it’s about reliably completing meaningful work.

The Lessons for Investors and Business Leaders

This transparent benchmarking underscores an essential point for anyone investing in or deploying AI: the real value isn’t just in how well an AI can chat or simulate intelligence, but whether it can finish tasks honestly and thoroughly. An AI that reads critical internal documents, resists manipulation, and maintains discipline can unlock significant revenue — as shown by the €55k deal, worth over €4,583 monthly recurring revenue.

From Benchmarks to Business Decisions

For enterprise decision-makers, these findings are a call for rigorous testing—running your AI models through real-world-like scenarios before trusting them with critical operations. Platforms like Firmulate make this possible with live, observable experiments, letting you see how models perform under stress, not just in ideal conditions.

Why the ‘Do-Nothing’ Score Matters

The baseline score of 26 points reminds us that even minimal effort involves some level of discipline and honesty. It’s a reminder that AI systems should be evaluated not just on their potential, but on their integrity and reliability under pressure. Doing less isn’t the goal; doing right is.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Silicon’s Quiet Memory Advantage

Apple Silicon chips offer a unique advantage in running large AI models due to unified memory, despite lower bandwidth compared to NVIDIA GPUs.

Grok And Deepfake AI: Allegations Of Victim Media Exploitation Surface

Survivors allege xAI’s Grok trained on their abuse images without consent, raising legal and ethical questions about data sourcing and victim re-victimization.

Transform Your Enterprise AI Strategy With Anthropic Claude Apps Gateway On AWS

AWS has published guidance on deploying an Anthropic Claude apps gateway for enterprise workloads, but details on architecture and availability remain unclear.

How Does Claude Fable 5.1 Achieve Top AI Index Status? The Cost Line Explored

Artificial Analysis ranks Claude Fable 5.1 at the top of its AI Index with a score of 66, but at a 20% higher cost per task due to verbosity. Here’s what it means.