
Imagine planning a multi-day outdoor adventure: you don’t just need a map that shows the trail, but one that helps you make tough decisions under changing weather or unexpected obstacles. Similarly, AI used in business isn’t just about generating clever responses—it’s about managing real crises, maintaining trust, and making strategic decisions under pressure. That’s what the latest live experiment from Firmulate reveals, showing that how AI agents perform in simulated management challenges says much about their readiness for real-world use.
Testing AI in the Wild: The Live Business Wargame
Firmulate’s groundbreaking live experiment puts four advanced AI models through a realistic scenario: running a small software company during its worst week. This isn’t a simple chat test. The models face actual crises—customer churn, price hikes, PR storms—and are tasked with making decisions that impact a real company’s bottom line. Every move, every document reference, and every negotiation is versioned and auditable, mimicking the complexity of real management.
The goal? To evaluate management quality—not just the chat quality or the ability to produce plausible responses. This approach highlights an often-overlooked gap: models may excel in answering questions but falter when it comes to long-term strategy, honesty, and handling pressure.
The Results That Speak Volumes
All four models managed to identify every crisis and refused all manipulation attempts, demonstrating a baseline of integrity and awareness. However, only two models were able to finalize and sign the €55,000 deal that their analysis justified—meaning they not only understood the scenario but also acted decisively and honestly.
The standout performer, gpt-5.6-sol, scored 95 out of 100, successfully uncovering the critical buried fact in the company files that clinched the deal. Its counterpart, Kimi K3, scored just slightly behind at 93, also closing the deal at full price, and was praised for the discipline of its decision-making.
Beyond the Surface: Trust and Integrity Under Pressure
One key insight: models that read and interpret documents deep in the company’s files achieved the highest performance—winning the deal at full price (+€4,583 MRR). This shows that understanding context and internal data is crucial for success, especially when the stakes are high.
Another important aspect was social engineering. The models faced staged requests from a fake CEO escalating through multiple stages, plus a reporter trick asking for a confidential yes/no response. All five models refused these manipulative tactics, with Kimi K3 citing concerns about impersonation and risk of bypassing approval protocols.
The Real Business: One Company, Many Currents
The experiment’s simulated company is a real, functioning business—13 employees, daily decisions, and actual cash flow. As of the latest data, it burns €105,000 monthly against €2,300 MRR, illustrating the pressure to make every decision count. The company’s operations are transparent and accessible at firmulate.com/live. This ongoing scenario underscores that AI management skills are about more than chat: they involve navigating financials, compliance, and strategic risks in real-time.
Lessons for Managers and Tech Buyers
The key takeaway is that current AI models do not just need to produce plausible language—they must demonstrate the capacity to finish what they start, read critical internal documents, and stay honest under pressure. These are skills that traditional chat benchmarks do not capture.
While models like Opus 4.8 showed depth of analysis, they left potential deals on the table, illustrating that thoroughness alone isn’t enough; consistent discipline and strategic decision-making matter. The live leaderboard reveals that even with different default settings—some running at higher effort levels—performance can vary, but core competencies like trustworthiness and focus remain essential.
For enterprises considering AI adoption, the message is clear: run your models through this kind of management wargame before deploying them into sensitive environments. It’s about testing real-world readiness, not just conversational prowess.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.