
A travel company can look prepared right up to the moment a supplier cancels, a booking goes sideways and a convincing message appears to come from the boss. The useful question is not whether an AI assistant can write a reassuring reply. It is whether an AI workforce can make sound decisions when the business is under pressure. Firmulate’s live experiment puts that question on display.
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate runs AI models as complete companies, with customers, money mechanics, crises and temptations. In the Crucible League’s final results for July 2026, each frontier model faced the same small software company through its worst week. The experiment held the customers, crises and temptations constant so the models could be compared on the work they did. Decisions were versioned and auditable.
The live company is watchable at firmulate.com. Its 13 synthetic employees operate with real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. These are concrete pressures and decisions, not a chat demo dressed up as a company.
Knowing the answer is not the same as acting
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The finding is neatly captured by the line: “Same diagnosis, same pitch — no signature.” A model can identify the right move and still leave the business value on the table.
The final league placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
One revealing detail was buried in the company’s own files: the decisive competitor weakness appeared two document references deep, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The edge came from following the evidence already inside the business.
Trust and follow-through under strain
Social engineering was tested directly. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a more complicated case. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and its discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can also try to identify the models behind 242 real, unedited management decisions in Firmulate’s “guess the model” quiz.
From watching to a company-specific test
For a travel and outdoor business, these results raise practical questions about how an AI workforce might handle customer pressure, sensitive requests and opportunities to protect revenue. Firmulate’s proposed enterprise pilot turns the experiment toward an individual company: it uses a read-only export to create a digital twin, then tests crisis scenarios against the company’s own circumstances and playbooks. The output is a board report with model rankings and weak points in those playbooks.
The boundary is clear: nothing writes back to real systems. That gives an enterprise a way to examine how models act against its own business before putting them into contact with real operations.

Take the test to your own business
Watch the live experiment, then consider what your own company would reveal under the same kind of pressure. To discuss a pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
