
Imagine running your health and wellness business with an AI so nuanced that it can make management decisions, prioritize tasks, and even refuse to cut corners—just like a seasoned human leader. What if these AI models each had distinct personalities shaping their choices? Welcome to the frontier of AI management—where personality, integrity, and decision-making are measurable and testable in real business scenarios.
The Experiment: Putting AI Models to the Test in a Business Crisis
At the heart of this groundbreaking experiment, four state-of-the-art AI models were tasked with managing a small software company going through its worst week. The setup was real: same customers, identical crises, and temptations to manipulate the system. Every decision was recorded, versioned, and made auditable, creating a transparent view into how each AI would behave under pressure.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management Personalities Through Performance Scores
The models entered a league table based on their ability to recognize crises, refuse manipulative offers, and close profitable deals honestly. The scores ranged from 73 to 95—indicating not just technical performance but management-like discipline and integrity.
Results: Different Personalities, Same Crises, Divergent Outcomes
- gpt-5.6-sol scored the highest with 95 points, successfully identifying a hidden critical document reference in the company’s files that clinched a €55,000 deal—full price, with full integrity.
- Kimi K3, a newcomer, scored 93 and also closed the big deal, demonstrating disciplined decision-making. Its reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
- Sonnet 5 scored 88 and was able to close the deal as well, but with some lapses in process discipline—a sign of a management style that is effective but less cautious.
- Fable 5 scored 77, also closing the deal but with weaker process adherence, leaving some opportunities on the table. Even the most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, slipped into hesitation, leaving a close opportunity unseized and discipline slackening.
The critical insight? All models recognized crises and refused manipulative tactics, including staged social engineering attempts like fake CEO messages and reporter tricks. The models demonstrated integrity across the board, but their management styles differed—some more thorough, others more aggressive or cautious.
The Hidden Weakness: Reading Deeper in Company Files
The real game-changer was the ability to access and interpret company documentation. The models that looked two document references deep into the company’s files were able to identify the critical information leading to the full-price deal. Those that failed to probe thoroughly left money on the table, highlighting a management personality that values diligence and depth.
Implications for Business and Wellness Enterprises
This experiment is not just about managing a software company; it offers a mirror for any business, including health and wellness brands, about how AI can embody different management personas. Will your AI prioritize thoroughness over speed? Integrity over shortcuts? The answers impact trust, profitability, and reputation—especially in sectors where authenticity and honesty matter as much as service quality.
Try It Yourself: The Firmulate Management Decision Quiz
Curious about how your organization’s AI might behave? Try the interactive quiz at firmulate.com/quiz.html to see which management personality your AI models resemble. It’s based on real, unedited decisions from actual business scenarios—no fiction, just facts you can learn from.
Experience the Live Business Simulation
For a hands-on experience, businesses can run their own management wargames using a read-only version of this experiment. No system writes back, but you get to see your AI’s decision-making style in action. Visit firmulate.com/pilot.html to learn more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html