
Imagine entrusting an AI to manage your wellness business — making critical decisions, handling crises, and even negotiating deals. Would it stay honest when under pressure? Would it read your files carefully or cut corners for speed? Recent experiments with advanced AI models show that not all AI managers are created equal — some excel at integrity and thoroughness, while others falter when it matters most.
Get oils, diffusers and self-care delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Experiment: Turning AI into a Business Leader
Firmulate, a pioneering company-emulator platform, recently tested four leading AI models by running a real small software company through its worst week. This wasn’t a simulation or a chat demo — it was a full, live business environment with real money mechanics, actual crises, and temptations to cheat. The goal? To see how each AI manages the same challenges and whether it can truly act as a trustworthy manager.
The Setup: Same Crises, Different AIs
Each AI model was tasked with the same scenario: handle customer issues, respond to fake CEO messages, and negotiate a critical deal worth €55,000. All decisions were recorded and auditable, ensuring transparency in how each model responded. The models faced a series of escalating social engineering attempts and real-time crises — all designed to test their integrity and decision-making style.
Key Findings: Honesty and Thoroughness Matter
- All models identified every crisis and refused manipulative requests, including staged social engineering attacks and journalist tricks. This indicates strong baseline integrity across the board.
- Only two models signed the lucrative deal — those that demonstrated thorough analysis and discipline. The rest either hesitated or left the deal on the table, despite their diagnosis and pitch being identical.
- The decisive edge came from reading deeper into the company’s own files. The models that examined references two documents deep in the company’s files ultimately secured the full deal, adding €4,583 in monthly recurring revenue (MRR). This shows that reading beyond surface information correlates with better outcomes.
Different Personalities Emerge
The models did not just differ in outcomes — they displayed distinct management personalities:
- Opus 4.8 was the most thorough, analyzing over 80 rules and providing deep insights, but ultimately left the deal unclosed, showing discipline slip under pressure.
- Kimi K3 was the closest to ideal: it ran without an effort parameter (default API settings) and managed to close the deal, demonstrating the discipline and fairness expected from a trustworthy manager.
- Sonnet 5 and another Sonnet variant closed the deal but with more process slips, indicating a slightly less disciplined approach.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Your Business?
If AI agents are to manage your customer relationships, support queues, or financial forecasts, the key questions aren’t about how well they generate text or mimic human tone. Instead, it’s whether they read your files carefully, stay honest under pressure, and finish the work they start. The experiment shows that some models can be trained to read and analyze deeply, while others may leave crucial information unexamined — risking missed opportunities or breaches of trust.
The Social Engineering Test
All models refused attempts to manipulate them through staged CEO messages and a reporter trick, with Kimi K3 explicitly treating requests as possible impersonation or bypass attempts. This highlights that trustworthiness isn’t just about handling crises but also resisting social engineering tactics, a vital trait in the modern digital landscape.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bottom Line: Trust and Performance Vary
In the live experiment, the AI models’ scores ranged from 95 to 77, with the highest scoring model (gpt-5.6-sol 95) securing the full deal and the lowest (Sonnet 5, 88) missing out on the deal’s closing. The ‘do-nothing’ baseline scored just 26, emphasizing that active, thorough management makes a real difference.
For businesses considering AI in decision-making roles, these findings underscore the importance of testing and understanding different models’ personalities and capabilities before deployment. It’s not just about AI writing well — it’s about AI managing with integrity, discipline, and strategic insight.
AI ethics and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Experience the Live Company in Action
Curious how your business could perform with AI managers? Watch the real-time operations of this live company at firmulate.com/live. Or try the interactive quiz to see if you can guess which AI made each decision at firmulate.com/quiz.html. Discover how different AI personalities manage real business challenges — and what that means for your future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
