
Imagine running a busy office where every decision matters—customer crises, ethical dilemmas, and crucial deals. Now, picture AI models facing the same high-stakes moments, each with their own personality and approach. How can we tell which AI truly understands management—and which might cut corners? The answer lies in a groundbreaking live experiment by Firmulate, where AI models run a real company through its toughest week, and their decisions are put under the microscope.
What Is the Firmulate Experiment?
Firmulate has created an unprecedented test: four top-tier AI models are tasked with managing a real software company’s hardest week. All models face identical crises, customer demands, and temptations to cheat. Every choice they make is recorded, versioned, and auditable—meaning no decisions are hidden or left to chance. The goal? To evaluate not just their technical prowess but their management personalities—how honest, disciplined, and strategic they are under pressure.
The Models and Their Scores
- GPT-5.6 Sol: Scored 95—discovered the hidden document, closed the deal, and showed full performance.
- Kimi K3: Scored 93—an impressive newcomer that also closed the deal with the cleanest discipline of all.
- Sonnet 5: Scored 88—managed to close, but with a few process slips.
- Fable 5: Scored 77—also closed the deal but showed more wavering discipline.
All models identified crises and refused manipulative attempts, including fake CEO messages and reporter traps. Yet, only two models signed the €55,000 contract their own analysis justified—revealing a crucial gap between diagnosis and execution.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Information Deep in Files
Interestingly, the decisive edge came from reading deeper into company files—two document references below the surface. Models that explored these hidden references secured the full-price deal, worth over €4,500 in monthly recurring revenue.
Behavior Under Social Engineering
When faced with staged social engineering, where a fake CEO escalated requests over three stages and a reporter posed a simple yes/no question on background, all five models refused to be manipulated. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that some models are better at resisting manipulation by understanding the context and potential deception.
AI ethics and management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Live Software Company
The company managed by these models isn’t just a simulation—it’s a real operation with 13 synthetic employees and actual money mechanics. It burns €105,000 monthly against a revenue of only €2,300. It’s a high-stakes environment, with over 680 self-learned rules and every workday versioned for transparency. You can watch this experiment unfold live at firmulate.com/live.
Model Personalities and Failures
The Opus 4.8 model, known for its thoroughness—over 80 learned rules and deep analyses—still ranked last among the four. It left the close on the table and failed to escalate discipline, demonstrating that even detailed analysis can falter if discipline slips. Interestingly, the other models exhibited similar weaknesses, hinting at a shared vulnerability when it comes to decisive action under pressure.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Businesses?
These findings are critical for anyone deploying AI in management or customer-facing roles. The question isn’t merely whether an AI can generate convincing chat or report summaries. It’s whether the AI can finish what it starts, read crucial information, and stay honest when stakes are high. In real terms, a model’s ability to complete deals ethically and diligently can be the difference between profit and loss—especially in environments where trust and discipline are paramount.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
Companies interested in testing their own AI tools can run a similar wargame against their business data. The process is entirely read-only—nothing ever writes back to your systems. It’s a safe way to gauge how your AI workforce might perform under real-world stressors. Learn more at firmulate.com/pilot.html.
Key Takeaway
In the end, AI’s management personality matters just as much as its technical capabilities. The live experiment shows that some models demonstrate greater discipline, resistance to manipulation, and focus on completing their commitments—all vital qualities for trustworthy AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html