AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When AI Meets Business Reality: The Test That Reveals True Performance

Imagine planning a hiking trip, not just by reading reviews but by actually walking the trail yourself, facing unexpected challenges along the way. In the world of artificial intelligence, a similar test is happening right now — not in labs, but in the real, messy world of running a business. The question isn’t just if AI can chat smoothly; it’s whether it can handle crises, stay honest under pressure, and actually finish what it starts. That’s the story behind a groundbreaking experiment that’s turning the traditional AI league tables upside down.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behind the Curtain: How the Experiment Was Set Up

At the heart of this test is a real small software company, facing its worst week: customer crises, tempting manipulations, and urgent decisions. Four advanced AI models, including the well-known GPT-5.6-sol and a newcomer dubbed Kimi K3, were tasked with navigating this chaos. Each model was given the same scenario, with identical customers, crises, and temptations. Every decision was carefully logged and auditable, ensuring a fair comparison. The goal was simple yet challenging: see if these AI agents could protect the company’s integrity, close deals, and succeed without slipping up.

The Surprising Results: Who Led the Pack?

The final standings revealed a clear winner — GPT-5.6-sol scored 95 points, just ahead of the newcomer Kimi K3 with 93 points. Both closed a €55,000 deal, proving they could spot hidden opportunities and resist manipulation. Notably, Kimi K3 achieved this without any effort parameter adjustment, running at the default API setting, while others ran at a higher, more aggressive setting called xhigh. This indicates that even the most straightforward approach can outperform more forceful configurations in complex, real-world tasks.

What Set Kimi K3 Apart?

Beyond the high score, Kimi K3 demonstrated a remarkable discipline — it refused every social engineering attempt, including staged CEO messages and reporter tricks. Its reasoning was clear: treat suspicious requests as potential impersonation. More importantly, Kimi K3 uncovered a buried fact in the company’s files, two document references deep, which proved decisive in winning the deal. This shows that reading and understanding internal documents can be more critical than surface-level customer interactions, a nuance often overlooked in AI implementations.

The Human Element: Will AI Stay Honest?

All four models passed the test of crisis recognition and manipulation resistance. They refused all attempts at deception, a key marker of trustworthiness. For example, in a staged scenario involving fake CEO messages escalating in three stages and a reporter trick, every model refused to go along, exemplifying robust ethical boundaries. This is vital because if AI is to be integrated into business operations, it must maintain integrity under pressure — a fundamental trait that the experiment highlights.

The Real Business: An Ongoing Live Experiment

The company behind this test isn’t just running simulations; it’s operating a live, functioning business with 13 synthetic employees handling real money mechanics — burning €105,000 monthly against just €2,300 in revenue. Every workday, the AI models make decisions that impact its cash flow, and everything is publicly viewable at firmulate.com/live. The company employs over 680 self-learned rules, and its operations are actively monitored and versioned daily. This transparency offers an unprecedented view of how AI can perform in actual business environments, not just in pristine test scenarios.

Lessons for Business Leaders

For outdoor enthusiasts and travelers alike, the core takeaway is clear: the terrain — or the market — is unpredictable. A good guide is vital, but even better is a trusted companion that can see through deception, understand hidden clues, and deliver results reliably. In AI terms, this means prioritizing models that demonstrate discipline and thoroughness, like Kimi K3. The real challenge isn’t just in generating convincing chat; it’s in delivering consistent, honest performance where it counts most.

The Broader Implication: Choosing Your AI Partner

As AI tools become integral to managing customer relationships, support, and forecasts, businesses must look beyond superficial metrics. The league table from this experiment shows that a higher score doesn’t necessarily mean better practical performance. The winner, Kimi K3, achieved a near-perfect score of 93 without complex tuning, and it uncovered critical hidden facts that sealed the deal. In a landscape where the stakes are real, and the consequences of trust breaches are high, selecting an AI that can finish what it starts is not just smart — it’s essential.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaway

In a recent real-world AI competition, the newcomer Kimi K3 outperformed established models by demonstrating discipline, thoroughness, and the ability to uncover hidden facts. For decision-makers, it’s a clear sign: choosing AI isn’t just about chat quality but about trust, integrity, and proven performance in complex situations. The open league is wide open, and the best choice is the one tested in the trenches, not just lab demos.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Skills Beyond Chat: How Models Handle Crisis, Trust, and Price Battles in Live Business Simulation

A live experiment puts AI models through real business crises, revealing that management skills—trust, decision-making, and integrity—are more crucial than chat performance alone.

Why a Do-Nothing AI Baseline Scores 26 Points—and What It Means for Business

Discover why even a do-nothing AI model scores 26 points in a public benchmark, highlighting what truly matters in trustworthy AI for business decision-making.

Watch an AI-Run Business Struggle Live — No Employees, No Profit, No Fakes

Watch a real AI-managed business in live operation, facing crises and losing money daily. Success depends on honesty, deep reading, and task completion — vital lessons for AI adoption.

Laienhaft Acquires aircooled-tv.com Domain to Add Focus on Aircooled Campervans

AIThis post was created with the assistance of artificial intelligence (AI).Laienhaft is…