
Imagine owning a kitchen gadget that not only cooks your meals but also manages your business—without any human employees, running daily on artificial intelligence. While that might sound like science fiction, a real experiment is happening right now, revealing what AI can and cannot do in the high-stakes world of company management.
The Bold Experiment: An AI-Managed Company in Action
At the heart of this daring venture is a small software company run entirely by AI models, with 13 synthetic employees and real money mechanics. This company is not just a simulation—it’s a live, functioning entity, available for anyone to watch at firmulate.com/live. Every workday, its decisions are recorded, versioned, and scrutinized, offering a rare window into how artificial intelligence handles real crises, temptations, and business decisions in real time.
As an affiliate, we earn on qualifying purchases.
How the AI Companies Are Tested
The experiment pits four advanced AI models—each with its own strengths and weaknesses—against the same week of chaos. These models include GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8, which have been given the same set of challenges, customers, and crises to navigate. Their decisions are fully auditable, so analysts can trace every move and assess performance objectively.
One of the core tests involves whether these models can spot hidden opportunities buried in company files—a task critical to closing deals and securing business. In this scenario, the models that read and analyze deeper documents outperformed their peers, winning an €4,583 monthly recurring revenue deal at full price, simply by uncovering a buried reference that others missed.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Ethical and Trust Challenges Under Scrutiny
The experiment also tests how models respond to social engineering attempts—fake CEO messages, staged reporter inquiries, and manipulative requests. All four AI models refused to be manipulated, with Kimi K3 explicitly treating suspicious requests as potential impersonation or approval bypasses. This indicates that, at least in these scenarios, the models maintain integrity and resist shortcuts designed to exploit them.
As an affiliate, we earn on qualifying purchases.
The Actual Company: An Ongoing Battle for Survival
The live setup is no idle simulation; it’s a real machine struggling against a €105,000 monthly burn rate against a meager €2,300 in monthly recurring revenue. The company’s cash countdown is public, and each workday’s decisions are recorded and analyzed, providing unprecedented transparency into AI’s operational limits in a business context.
AI ethical decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights From the Results
The models’ performance varied. Opus 4.8, despite its thorough analyses and 80+ learned rules, was the last to close a deal, leaving opportunities on the table and slipping in discipline. Kimi K3, running without an effort parameter (the default API setting), managed to sign the deal—demonstrating the importance of configuration choices in AI performance.
In the end, the top performer was GPT-5.6-sol, which uncovered the hidden deal opportunity and secured the revenue, effectively completing the company’s objectives. The scores—95 for GPT-5.6-sol, 93 for Kimi K3, 88 for Sonnet 5, and 77 for Opus 4.8—reflect their effectiveness and discipline under pressure.
Why This Matters for Your Business
This experiment is a powerful reminder that in AI-driven management, the real questions are not about conversational chat quality but about reliability, honesty, and the ability to follow through on crucial tasks. Will your AI support system read your files properly? Will it stay honest when tempted or pressured? Will it finish what it starts, or leave opportunities on the table?
What You Can Learn
Understanding the performance gaps and strengths of different AI models can help businesses make smarter decisions about deploying AI solutions—whether for customer management, support, or operational decision-making. The experiment’s results underscore that AI can identify hidden opportunities and resist manipulation, but configuration and discipline matter greatly.
Explore and Observe for Yourself
Visit firmulate.com/live to watch the live experiment in action, read detailed performance scores, and see real decisions made by each model. For more insights into what AI can do in a business setting—and how it might impact your operations—check out the quotes page for in-depth analysis.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html