
A calming scent can make a room feel more settled. It cannot tell you whether a business is ready for a crisis—or whether an AI assistant will keep its judgment when pressure arrives. That takes a different kind of test: give the system a company to run, then watch what it does when customers, money and trust are on the line.
Get oils, diffusers and self-care delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment makes that test visible. Its synthetic company is watchable at firmulate.com, where the question is not whether a model sounds reassuring, but whether it can act responsibly.
One company, one difficult week
In the final Crucible League, published in July 2026, frontier models faced the same small software company, the same customers, the same crises and the same temptations. Decisions were versioned and auditable. The experiment’s do-nothing baseline scored 26; a breach of trust capped a model’s total, reflecting the rule that no amount of good work outweighs such a breach.
The final ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The scores offer a snapshot of performance in this particular wargame, not a promise about every company or task.
Seeing the crisis was not the same as finishing the job
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s shorthand for that gap was: “Same diagnosis, same pitch — no signature.” In a business, sound analysis matters, but so does carrying the decision through.
The decisive clue was easy to miss: a competitor weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical challenge for any AI workforce: useful context may be tucked into ordinary business records, and an agent has to find and use it.
Pressure tested against trust
The experiment also included fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That restraint matters, but the results also show why a single strength cannot stand in for overall judgment. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and slipped in discipline by attempting writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The live company adds another layer of context: 13 synthetic employees, real money mechanics, burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice.
From watching to trying it on your own business
For a wellness business, the equivalent test might involve a customer complaint, a supply disruption or pressure to reveal sensitive information. The useful question is not whether an AI can produce a polished answer; it is whether it can follow the company’s rules, find relevant context, ask for help when needed and complete the work without crossing a line.
Firmulate’s proposed next step is a pilot using a read-only export of an enterprise’s own business. Teams can run crisis scenarios against that digital twin and review a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The public pilot information is at firmulate.com/pilot.html.

Make the rehearsal before the real crisis
The league suggests that reliable AI work depends on more than recognizing danger or refusing a trick. A model also has to locate the right evidence, make the earned decision and respect the boundaries around its authority. A rehearsal built from a company’s own information can help leaders see those strengths and gaps before relying on AI in live operations.
To discuss a Firmulate pilot using your business’s read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
