
Preppers know the value of a plan that has been tested before the emergency. A checklist can look solid on paper and still fail when people face pressure, conflicting signals and a fast-moving crisis. Companies are beginning to ask the same question about AI: what happens when an agent has to make decisions during a genuinely bad week?
Get emergency and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts that question on public display. Its simulated company has synthetic employees and real money mechanics, and its AI participants face the same customers, crises and temptations. The point is to watch what they do, not just what they say they would do.
One company, five very different finishes
In the final Crucible League, published in July 2026, five models ran the same small software company through its worst week. The results ranged from 95 for gpt-5.6-sol to 73 for Opus 4.8, with Kimi K3 at 93, Sonnet 5 at 88 and Fable 5 at 77. A do-nothing baseline scored 26. The league’s rule is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
All the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding is a useful reminder for anyone planning for a breakdown: recognizing the emergency is not the same as carrying out the response. The experiment summed up the gap as “Same diagnosis, same pitch — no signature.”
The clue was already in the company’s files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It did not appear in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. In other words, the advantage came from finding and using relevant information that was already available, then following through.
That distinction matters in a business crisis. A team can spot the obvious warning and still miss the detail that changes the outcome. Firmulate’s experiment makes the decision trail watchable, with every decision versioned and auditable. Readers can also try to identify which model made real management decisions in a quiz built from 242 real, unedited decisions at Firmulate.
Pressure tests include trust and discipline
The manipulation test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For organizations weighing AI access to customers, support queues or forecasts, that kind of refusal is part of preparedness: a system under pressure must preserve boundaries as well as identify threats.
Refusal alone, however, did not guarantee a strong result. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models. The lesson is not that a detailed analysis guarantees a good outcome. It is that a playbook also has to make the next safe and useful action clear.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also reports that the live company has 13 synthetic employees, burns €105k a month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Those mechanics make the experiment concrete, while keeping its company simulated.
From watching to a company’s own drill
For enterprises, Firmulate offers a pilot that applies the wargame to a read-only export of their own business. The export can provide a basis for testing crisis scenarios against company-specific information and examining weak points in existing playbooks. The exercise produces a board report with model rankings and findings about those weak points. Nothing writes back to real systems.
That boundary is central to the idea: rehearse decisions against a copy of the business before an AI agent is trusted with live operations. The public experiment shows the kind of gap a drill can reveal—spotting a crisis, refusing a trap, finding evidence and still failing to close the loop. The pilot lets a company ask how its own models and procedures handle that sequence.

Run the drill before the stakes are real
Preparedness depends on more than having a plan. It depends on knowing whether people and systems can find the right information, resist manipulation and follow a safe path to action when pressure rises. Firmulate’s live experiment makes those behaviors visible; its enterprise pilot brings the exercise to a company’s own business data.
To discuss a pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
