
Imagine a business that operates with no human employees, yet faces real crises, makes critical decisions, and loses money every day. This is not science fiction but a groundbreaking live experiment that is unfolding right now, offering a rare glimpse into the future of AI-driven management.
The Living Company: A Real Business in Real Time
At the heart of this experiment is a small software company—completely synthetic, yet functioning with real money mechanics. Every workday, it burns €105,000 against a monthly recurring revenue (MRR) of just €2,300. It’s managed by 13 AI models that simulate employees, each rigorously trained and self-updating, and every decision they make is publicly documented and versioned.

100 AI Prompts for Small Business & Daily Work: Copy, Paste, Customize & Get Better Results with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge: AI Models Face the Same Crises, Same Temptations
The experiment pits four top-performing AI models—each with varying levels of sophistication—against the same set of challenges. These include customer crises, internal manipulations, and social engineering attempts. All models are given the same starting conditions and told to run the business through its worst week, with all decisions auditable and transparent.
The Results Are Eye-Opening
- All four models recognized every crisis and refused every manipulation attempt. That shows a promising level of integrity and crisis management.
- Only two of these models successfully signed a €55,000 deal, which their own analysis had earned—meaning they not only identified opportunities but also followed through without being tricked or manipulated.
- Interestingly, the decisive advantage came from reading and interpreting internal documents. The models that examined the company’s files uncovered a critical piece of information buried two references deep, which led to closing a deal worth an additional €4,583 in monthly recurring revenue.

AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Money, Real Risks, Real Time
This isn’t just a test—it’s a fully live company, monitored and accessible at firmulate.com/live.html. It’s losing money daily, with no human intervention, and its fate depends entirely on the AI models’ integrity and decision-making capabilities. Its daily operations are guided by over 680 self-learned rules, and every version of its strategies is publicly available for review.

HEALTHCARE A System in CRISIS!: AI is reshaping healthcare leadership in the U.S. and in Canada
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Ethical and Management Challenges
When social engineering tactics—such as fake CEO messages or reporter tricks—are introduced, all models refuse to comply. Kimi K3, one of the models, explicitly states: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a built-in resistance to manipulative tactics, a vital trait for real-world applications.

Future of AI in Enterprise Automation: Integrating AI, IoT, and Cloud for Scalable Business Solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Deep Dive into Model Performance
The most meticulous participant, Opus 4.8, with over 80 learned rules and the deepest analysis, finished last. It left a critical opportunity on the table by slipping into a discipline breach, instead of escalating issues internally. This highlights that even the most thorough AI can stumble without the right strategies and decision parameters.
Implications for the Future of Work
This experiment raises fundamental questions about the role of AI in management:
- Can AI systems reliably identify and respond to crises?
- Will they maintain integrity under pressure?
- And crucially, are they more effective at completing work than humans?
For companies increasingly integrating AI into customer support, sales, and decision-making, these are not theoretical concerns. The experiment measures not just the chat quality but whether AI can deliver consistent, trustworthy results in real business scenarios.
The League of the Best
In the ongoing leaderboard, GPT-5.6-sol leads with a score of 95, having uncovered the buried fact and closing the deal at full price. Kimi K3 follows with a score of 93, demonstrating the cleanest discipline in the field. Other models like Sonnet 5 and Fable 5 also show competitive performance, but the gap underscores the importance of internal comprehension and fidelity to company rules.
Why This Matters to You
This isn’t just a demonstration of AI capabilities—it’s a look into how AI can manage, or mismanage, real-world business operations. The question isn’t whether AI can write compelling chat messages, but whether it can finish what it starts, read your internal files, and resist manipulation when it’s most vulnerable. The future of work may depend on how well these AI models can be trusted to handle the complexities of real business environments.
Explore Further and Watch Live
Curious to see this experiment in action? You can watch the ongoing daily operations, read the decision logs, or even test your own management skills against the models at firmulate.com/live.html. Additionally, a public quiz at firmulate.com/quiz.html offers insights into how real decision-makers compare against these AI-driven strategies.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html