
Imagine a business that’s entirely run by artificial intelligence—no human employees, no managers—yet it’s struggling for survival every single day. Sounds like science fiction? Well, it’s happening right now, in real time, and you can watch every move unfold at firmulate.com/live.
The Live Experiment: A Company in Crisis
At the heart of this project is a small, software-based company controlled by four different AI models, each tasked with handling the same challenging week of work. The company’s core mechanics are real: it burns €105,000 each month while generating only €2,300 in monthly recurring revenue—an unsustainable situation. Yet, what makes this experiment extraordinary is that every decision, crisis response, and temptation to cheat or manipulate is publicly recorded and auditable.
How Do AI Models Perform in High-Stakes Management?
All four AI models faced identical crises—customer issues, internal dilemmas, and even social engineering attempts like fake CEO messages or reporter tricks. Remarkably, all four recognized every crisis and refused manipulation attempts. That’s a first: these AI agents demonstrate an impressive ability to stay honest under pressure.
However, when it came to closing deals based on the insights they uncovered, only two models managed to sign the €55,000 deal their own analysis had justified. The other two identified the opportunity but left the deal unexecuted, illustrating that detection alone isn’t enough—execution matters too.
AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness in the Company’s Files
A crucial detail emerged from the company’s own documents—information buried two references deep in their files, not immediately visible during crisis management. The models that read and understood the full context of these documents succeeded in closing the full-price deal, adding +€4,583 monthly recurring revenue. This underscores the importance of thorough information processing—something that even an AI can struggle with if not designed to dig deep enough.
Social Engineering and AI Integrity
In a staged social engineering test, fake messages from a supposed CEO and a reporter tricked some models into considering shortcuts or approvals bypassing proper channels. All five tested models refused to act impulsively, with Kimi K3 explicitly reasoning about potential impersonation risks. This indicates a promising level of ethical reasoning and caution in these AI agents.
AI ethics and security tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Building in Public: The Real-World Struggle
The experiment is not just about AI’s technical ability; it’s about transparency, discipline, and the limits of automated management. The company, called the ‘live company,’ operates with 13 synthetic employees and constantly updates its playbook—over 680 self-learned rules—to handle daily crises. Every workday is versioned, and the entire process is open for public observation.
The company is currently burning €105,000 a month against minimal income, with a countdown to running out of cash. Yet, despite the losses, the models show promise—especially in recognizing crises and maintaining ethical standards. The question is whether this approach can someday replace or augment human management, especially for small businesses wary of hiring full-time staff.
Lessons from the Deep Dive: Opus 4.8 and the Fight for Discipline
The most thorough model, Opus 4.8, analyzed over 80 learned rules but still left a deal on the table and slipped in discipline, showing that even the best AI has weaknesses. Interestingly, the same flaws appeared across all models, hinting at fundamental challenges in managing complex, real-world business situations solely through AI.
AI business automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Everyone
If AI agents are to someday manage your customer relationships, support queues, or forecasts, the key isn’t just how well they chat or mimic human conversation. It’s whether they finish what they start, read your documents thoroughly, and stay honest under pressure. These are the real measures of their usefulness—and the simple truth that’s visible in this experiment’s transparency.
Through watching this live trial, you see a glimpse of the future: AI as a management tool that is tested daily against crises, temptations, and ethical dilemmas. It’s not a perfect story—yet. But it’s an essential one, unfolding in real time for anyone interested in the future of work and automation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI crisis management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.