The AI Race Is Not Over After The Demo—Here’s The Leaderboard To Watch
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get smart everyday buys delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The AI competition remains fierce after a live management test, with GPT-5.6-SOL leading the leaderboard. The experiment emphasizes management quality and trust, not just responses, as detailed in the original analysis. Key results reveal strengths and gaps in AI decision-making.

The latest AI management benchmark, conducted by Firmulate, has concluded with GPT-5.6-SOL ranking first on the leaderboard, achieving a score of 95 out of 100. This live experiment tested AI models in a simulated company environment facing crises, trust challenges, and decision-making under management scenarios. The results underscore that management skills—such as investigation, communication, and trust—are now critical measures of AI performance, beyond traditional chat or coding benchmarks.

The experiment involved five AI models managing a small software company during its most difficult week, with real financial stakes and a strict trust policy. GPT-5.6-SOL outperformed competitors with a high score of 95, while Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. The models were tested on their ability to diagnose crises, communicate decisions, and uphold trust, particularly in scenarios involving manipulation attempts and safety concerns.

Despite strong performance in identifying crises and refusing manipulation, only two models managed to close a significant €55,000 deal based on their analysis. Notably, the winning model excelled at diagnosis but faltered in presenting the critical fact that would clinch the sale, buried deep in internal files. This highlights a key insight: an AI can sound informed yet fail to retrieve essential information, affecting real-world outcomes. The experiment also revealed that thoroughness and activity do not necessarily translate into effective management, as Opus 4.8, despite its detailed analysis, finished last due to lapses in escalation and decision execution.

At a glance
updateWhen: final results announced July 2026
The developmentThe AI management benchmark concluded with GPT-5.6-SOL at the top, demonstrating that AI leadership in real-world scenarios is still evolving.

Why Management Skills in AI Matter More Than Ever

The results demonstrate that AI’s value in business extends beyond generating plausible responses. Effective management requires the ability to prioritize, read organizational context, maintain trust, and execute decisions reliably. As AI models are increasingly integrated into operational roles, their capacity to handle real-world consequences becomes a critical measure of success. The experiment underscores that trust breaches and decision accuracy are more revealing than superficial performance metrics, emphasizing the need for evaluation frameworks that focus on management quality.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution of AI Benchmarks in Business Management

Traditional AI benchmarks have primarily focused on technical outputs like coding accuracy or conversational quality. However, recent efforts, including Firmulate’s live company test, aim to evaluate models in dynamic, high-stakes environments. This approach reflects a broader shift toward assessing AI’s ability to manage real-world complexities, including crises, trust, and organizational decision-making. The July 2026 leaderboard marks a significant milestone, showing that AI’s management capabilities are now a critical frontier in AI development and evaluation.

“Management quality, not chat quality, deserves its own category of AI evaluation.”

— Thorsten Meyer, founder of Firmulate

Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Long-Term AI Management Performance

It remains unclear how these models will perform over extended periods or in different organizational environments. The experiment focused on a single, high-pressure week, and results may vary with different scenarios or larger organizations. Additionally, the impact of model training, architecture, and operational parameters on management skills needs further exploration. As AI continues to evolve, understanding how these factors influence real-world management remains an open question.

Amazon

business AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Future evaluations will likely expand to longer-term scenarios, testing models’ ability to adapt and learn over time. Companies considering AI for management roles should prioritize assessing trust, decision accuracy, and escalation practices. Industry leaders are expected to refine benchmarks further, incorporating real-world feedback and diverse organizational contexts. The ongoing development aims to establish management quality as a standard measure for AI readiness in operational settings.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the leaderboard tell us about AI’s management skills?

The leaderboard shows that AI models are improving in diagnosing crises and refusing manipulation but still struggle with execution and closing deals, highlighting areas for further development.

Why is trust important in AI management performance?

Trust determines whether an AI can make decisions that uphold organizational integrity, especially under pressure or when facing manipulation attempts, making it a critical evaluation criterion.

Can current AI models replace human managers?

While models show promise in specific tasks, they still lack the comprehensive judgment, contextual understanding, and trustworthiness required for full managerial roles, especially over long periods.

What should companies consider before deploying AI in management roles?

Organizations should evaluate AI’s ability to read organizational context, prioritize tasks, escalate when necessary, and maintain trust—beyond just assessing response quality.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kimi K3’s Journey To #3 On VigilSAR’s Public AI Rankings

Kimi K3 by Moonshot debuts at #3 on VigilSAR’s public AI ranking for defense-ISR models, surpassing GPT and Gemini models in trustworthiness metrics.

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

Learn the strategies to make your AI stack kill-switch-proof amid US government shutdowns and export controls, emphasizing dependency mapping and open-weight models.

How Technology Is Improving Food Safety In The Restaurant Industry

New vision-model inspection tools are being tested to improve food safety checks, offering verifiable data for restaurant operations and quality assurance.

SK Telecom Pursues 15GW AI Data Center Buildout, Aiming To Become Asia’s AI Infrastructure Hub

SK Telecom announces plans to build a 15GW AI data center network, aiming to become Asia’s leading AI infrastructure provider, according to PR Newswire.