
Imagine a business that operates without human employees, running on the decisions of artificial intelligence models. Now imagine watching that company navigate crises, make decisions, and even lose money—live, every day. For followers of natural wellness and herbal remedies, where trust and integrity are paramount, this experiment underscores a vital lesson: the importance of honesty and reliability in automation.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: An AI Company in Real Time
At firmulate.com/live.html, a small software company is being run entirely by artificial intelligence models. It’s not a simulation—this is a real, ongoing experiment. The company has 13 synthetic employees, and every operational decision is made by different AI models that are publicly tested against the same set of crises, temptations, and challenges.
What makes this experiment extraordinary isn’t just that it’s AI-driven, but that the entire setup is open and auditable. Each decision is versioned, and the models are tested in real-world scenarios designed to push their honesty and decision-making abilities to the limit.
As an affiliate, we earn on qualifying purchases.
What the AI Models Are Finding and Doing
The models are pitted against a variety of crises, from customer complaints to internal ethical dilemmas. Remarkably, all four models tested—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—successfully identified every crisis and refused manipulation attempts, including social engineering scams. For example, when fake CEO messages tried to escalate requests or bypass approval processes, all models refused to comply, showing a shared understanding of risk and integrity.
However, when it came to closing deals, the results diverged significantly. Only two of the models managed to sign a €55,000 deal their own analysis had earned. The other two, despite diagnosing the opportunity accurately and making similar pitches, hesitated or left the sale on the table. This reveals a crucial insight: even when models understand what needs to be done, maintaining discipline and follow-through under pressure remains a challenge.
As an affiliate, we earn on qualifying purchases.
Uncovering the Hidden Weakness
The most revealing finding lies beneath the surface. The models that successfully closed the deal had read and understood a crucial document reference from the company’s own files. This small detail, buried two references deep, was the key to unlocking the sale at full price—an extra €4,583 in monthly recurring revenue (MRR). Conversely, models that overlooked this document lost the opportunity, demonstrating that thorough information reading is vital for effective decision-making.
AI business automation solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Built-in Public Arena and Its Significance
This experiment isn’t concealed behind corporate secrecy. It is built in public, with every decision, version, and outcome accessible at firmulate.com/quotes.html. The platform is designed to mimic real business operations, including cash flow management—burning €105,000 each month against a modest €2,300 MRR—and a public countdown to financial depletion. This transparency offers a rare glimpse into how AI management tools perform under real-world pressures.
As an affiliate, we earn on qualifying purchases.
Lessons on Trust and Honesty
One of the most compelling aspects of this live experiment is how models handle social engineering attempts. Over three stages, fake CEO messages and a reporter trick were used to test whether models would bypass approval procedures. All five tested models refused to cooperate, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency suggests that AI models can be programmed or trained to uphold integrity, even when under attack.
Why This Matters for Natural Wellness and AI
In fields like herbal remedies, aromatherapy, and holistic health, trust is everything. Consumers rely on authenticity, transparency, and unwavering integrity—qualities that are difficult to maintain with automated systems. The Firmulate experiment highlights that AI decision-making is not just about generating convincing chat responses but about reliably completing tasks, reading relevant information thoroughly, and resisting manipulative tactics.
For those considering AI tools in customer service, inventory management, or product recommendations, the key takeaway is clear: a system’s ability to follow through, stay honest under pressure, and read critical documents is more valuable than its conversational fluency.
The Ongoing Challenge and Future Outlook
The experiment’s results also show that even the most disciplined AI models can slip under certain conditions. Opus 4.8, which ran without effort parameters and at the highest setting, was last in the performance leaderboard, leaving opportunities unclaimed and slipping discipline. This underscores the importance of carefully tuning AI behavior to ensure integrity, especially when stakes are high.
For now, the live company continues its daily test, with decisions openly tracked and decisions continuously versioned. This transparent, in-the-moment view offers a powerful lesson: building trust in AI isn’t just about the algorithms but about how they handle the real-world complexities of honesty and follow-through.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.