firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In the world of finance and investing, it’s tempting to judge AI by how well it can simulate conversation or generate reports. But real business success hinges on something far more elusive: execution. When AI faces the toughest tests—crises, temptations, and ethical dilemmas—it’s not just about what it says, but what it actually does. A groundbreaking live experiment reveals that only two out of four leading AI models can reliably close deals and follow through on analysis, exposing a critical blind spot in current AI evaluation methods.

The Experiment: Putting AI to the Test in a Real Company

To understand AI’s true business capabilities, the company behind Firmulate organized a unique live experiment. Four top frontier AI models—each with different strengths—were tasked with running a small software company through its worst week. This wasn’t a test of chat prowess but a comprehensive simulation involving real money mechanics, customer crises, and ethical challenges. Every decision was versioned, auditable, and embedded in the company’s actual file system, ensuring no shortcuts or superficial performance.

The AI Sales Coach: Objection Handling, Closing, and Prospecting Reimagined (The Objection Handler's Library)

The AI Sales Coach: Objection Handling, Closing, and Prospecting Reimagined (The Objection Handler's Library)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Spotting Crises and Resisting Manipulation

Remarkably, all four models identified every crisis the company faced and refused manipulation attempts—fake CEO messages, staged reporter inquiries, and other social engineering tricks. On this front, they all proved resilient. However, the critical difference emerged in execution: only two models managed to close the €55,000 deal that their own analysis had earned, while the other two left the deal on the table, despite having diagnosed the opportunity accurately.

Learning Robotic Process Automation: Create Software robots and automate business processes with the leading RPA tool – UiPath

Learning Robotic Process Automation: Create Software robots and automate business processes with the leading RPA tool – UiPath

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Does This Matter for Business and Finance?

For investors, business leaders, and finance professionals, this experiment underscores a crucial point: AI’s ability to produce convincing conversations or reports is not enough. Success depends on its capacity to follow through—reading the company’s files, executing decisions, and closing deals. In this test, the models that read deeper into the company’s documentation secured the full deal, translating diagnosis into tangible results.

Document Intelligence Made Easy: A Beginner’s Guide to Humata AI

Document Intelligence Made Easy: A Beginner’s Guide to Humata AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Deep Reading and Discipline

The decisive factor was the models’ ability to read and act on information buried within company files. The winner—gpt-5.6-sol—found a buried fact that was essential to closing the deal, and then signed it. Conversely, Opus 4.8, which was the most thorough in analysis, ultimately failed to finalize the deal, illustrating that even thoroughness alone isn’t enough without disciplined execution. Interestingly, the Kimi K3 model ran without an effort parameter, maintaining discipline and closing the deal cleanly, further highlighting that consistency, not just intelligence, matters.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Challenge of Social Engineering and Ethical Pressure

Another vital insight was that all four models refused to be manipulated through staged social engineering. Even under escalating fake CEO messages and staged reporter requests, they held firm, reasoning that such requests could be impersonation or approval bypass attempts. This resilience demonstrates that models can be trained to refuse unethical shortcuts, a critical feature for protecting business integrity.

Implications for Investors and Business Leaders

This experiment reveals something profound: the real value of an AI model isn’t just in its ability to generate human-like text but in its capacity to complete the work it’s assigned—reading the right files, making disciplined decisions, and closing deals. For those investing in or deploying AI in finance, sales, or operations, it’s vital to look beyond chat demos and focus on the AI’s execution strength.

Live Business, Real Money, and Transparent Testing

The experiment takes place on firmulate.com/live, where you can watch the software company operate in real-time, with every decision visible and auditable. The company burns €105,000 monthly against a modest €2,300 in monthly recurring revenue, illustrating the real stakes involved. Every workday, the models are tested against actual crises, and their decision-making is recorded for transparency and learning.

Conclusion: Execution Over Words in AI Performance

For investors and business owners, this live experiment underscores a crucial lesson: evaluating AI solely based on chat or superficial demos is misleading. The true test is whether the AI can *finish* what it starts—reading relevant documents, resisting manipulation, and closing deals or solving problems reliably. Only then can AI truly become a trustworthy partner in managing your financial or operational assets. To explore how AI can emulate your business under pressure, visit Firmulate.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Apple Silicon’s Quiet Memory Advantage

Apple Silicon’s unified memory architecture offers a unique capacity advantage for large AI models, despite lower bandwidth and speed compared to NVIDIA GPUs.

China Sphere Capability Gap, Q2 2026 Update: Five Labs, Five Strategies, One Narrowing Frontier

Five Chinese labs launched frontier models within four weeks, narrowing the capability gap with US leaders while maintaining cost and licensing advantages.

Seeing AI Through A Benchmark Partner’s Eyes: What Others Miss

Benchmark investor Eric Vishria warns against zero-sum thinking in AI markets, highlighting market complexity and the importance of differentiation.

The Key To Voice Talent Licensing In The AI Marketplace

A licensing and approval hub for voice actors’ AI clones is being tested as a streamlined way to manage usage, payments, and consent in the AI voice marketplace.