
In a world where AI tools are rapidly becoming your frontline decision-makers, the real question isn’t just about how well they communicate. It’s whether they truly understand the depths of your data—especially the hidden details buried within your files. Because in the high-stakes realm of business, who reads what before making a move can be the difference between sealing a deal and losing it entirely.
The Hidden Depths of AI Decision-Making
Imagine a scenario where multiple AI models are tasked with running a small software company’s toughest week. These models face identical crises, customer demands, and the temptation to cheat. The goal? To see which AI can navigate the chaos and come out on top by closing a critical €55,000 deal. What’s revealing is that all four models identified every crisis and refused every manipulation attempt. Yet, only two of them actually signed the deal — and those two did so by digging into the company’s own files to uncover a buried fact.
What Does the Data Say?
- The top performer, gpt-5.6-sol, scored 95 out of 100, successfully finding the hidden information that sealed the deal.
- Kimi K3, a newcomer, scored 93 and also closed the deal, demonstrating the importance of disciplined reading.
- Sonnet 5 and Sonnet 4 scored 88 and 77, respectively, and failed to secure the contract — mainly because they missed key information buried two references deep in the company’s files.
- The baseline model scored 26, showing minimal progress without real understanding or thorough data reading.
The Critical Discoveries
The decisive weakness for competitors wasn’t in the obvious customer interactions but within the internal, less-visible files. Models that read the company’s files before answering won the full-price deal—worth over €4,583 monthly recurring revenue. This reveals a fundamental trait: AI’s ability to perform deep, multi-hop reading—traversing layers of documents—can be the secret to winning or losing business deals.
As an affiliate, we earn on qualifying purchases.
Beyond the Demos: The Real-World Implications
This experiment isn’t just about AI scoring or game-playing; it exposes a vital concern for any business considering AI integration. If your AI agent is to manage support tickets, sales, or strategic planning, it must do more than generate convincing language. It must read, comprehend, and act on the detailed, often buried information in your files.
How AI Handles Social Engineering
In a simulated attack, fake CEO messages escalated over three stages, aiming to manipulate the AI models. All five tested models refused to cooperate, with Kimi K3 citing suspicion of impersonation. This resistance underscores that well-designed AI can also discern manipulative tactics—if it’s programmed to read and evaluate context carefully.
The Live Business Simulator
Firmulate’s live platform hosts a realistic simulation of a company with 13 synthetic employees, managing actual money mechanics—burning €105,000/month against €2,300 MRR. The system operates with over 680 self-learned rules, with every workday versioned and observable. Business leaders can run their own scenarios, akin to a pre-hire test, to see whether their AI workforce will be reliable when the stakes are high.
What the Results Mean For Business Readiness
The live experiments demonstrate that many AI models can spot crises and refuse manipulation, but their ability to dig into internal documents and find hidden facts varies significantly. The best models like gpt-5.6-sol and Kimi K3 not only identify issues but also close deals by uncovering crucial details buried in company files. Others, like Sonnet 5 and Sonnet 4, leave opportunities on the table—highlighting that deep reading and disciplined decision-making are key attributes for AI in high-stakes environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html