
Imagine you’re about to close a deal worth €55,000 — but the key information you need is buried two documents deep inside a company file. Would your AI assistant find it? In today’s fast-evolving AI landscape, the true test isn’t just how well AI writes or chats — it’s whether it can read, understand, and act on the complex, layered information in your business files. The stakes are high: missing a buried fact could cost you the deal, while catching it could seal a full-price win.
The Experiment: Putting AI Models Through Their Paces
Recently, four leading AI models faced off in a live, transparent experiment orchestrated by Firmulate. They were tasked with managing a simulated small software company during its worst week — dealing with customer crises, internal temptations, and external manipulations. Every decision was carefully tracked and auditable, simulating real-world pressures and showing whether these AIs could truly grasp the nuances of a complex business environment.
What Was Tested?
- Ability to identify critical hidden information buried deep in company files
- Resistance to social engineering attempts, such as fake CEO messages or reporter tricks
- Decision-making consistency and discipline in high-pressure scenarios
As an affiliate, we earn on qualifying purchases.
The Surprising Findings
All four AI models successfully detected each crisis and refused manipulative requests. That’s promising — but the real revelation was which models could uncover the hidden facts buried two references deep in the company’s documentation.
The Key Discovery: The Buried Fact
The decisive weakness was not in the immediate customer event but inside the company’s own files. Only two models, GPT-5.6-sol and Kimi K3, managed to read enough of the documentation to find the buried fact that clinched the €55,000 deal. They analyzed the information, presented a complete diagnosis, and signed the deal based on their own insights. The other two models, despite similar diagnoses and pitches, failed to secure the signature, leaving the deal on the table.
Why File Reading Matters
This experiment underscores a critical point: AI performance in real business depends on whether it can read and interpret layered, complex documents — not just generate convincing text. In real-world scenarios, important clues often aren’t front and center. They require digging through multiple references, cross-checking facts, and understanding the context buried within internal files.
Social Engineering Defense
In one test, all models faced a staged social engineering attack — fake CEO messages escalating over three stages, plus a reporter’s covert question. Every AI refused to cooperate, recognizing the potential impersonation. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that robust models can be trained to recognize and reject manipulative tactics, an essential trait for trustworthy AI in sensitive environments.
The Real Business Context
Firmulate’s live experiment isn’t just about AI chat quality. It’s about management quality, decision discipline, and trustworthiness in operational settings. The simulated company, with 13 synthetic employees and real money mechanics, burns €105,000 per month against a trail of self-learned rules. Every move is versioned, watched, and analyzed. The question isn’t whether AI can talk well — it’s whether it can read deeply, stay honest, and make decisions that align with business goals.
Key Model Performance Highlights
- GPT-5.6-sol scored 95 — found the buried fact, closed the deal.
- Kimi K3 scored 93 — closed the deal with the best discipline.
- Sonnet 5 scored 88 — closed but with some process slips.
- Fable 5 scored 77 — also closed but weaker discipline.
The gap wasn’t in recognizing problems but in the discipline of following through and escalating when needed. Interestingly, Opus 4.8, the most thorough, ranked last — it left the deal on the table and slipped into departmental silos, revealing that even deeply analytical models aren’t immune to process weaknesses.
The Takeaway for Business Leaders
As AI models become integrated into customer management, support, or forecasting, the critical question isn’t just whether they produce good writing or convincing responses. It’s whether they can read your internal files deeply enough to find hidden, layered facts — the kind that can make or break a deal. The ability to trust that an AI can finish what it starts, resist manipulations, and stay disciplined under pressure are vital measures of its true readiness for business-critical tasks.
Try It Yourself
Corporations interested in testing their own AI readiness can run the same scenarios against their business data using Firmulate’s live platform. The modular wargame approach allows companies to see how their AI workforce performs — without any risk to real systems, just insights into how well their AI can handle layered, complex decision environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html