AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Preppers know the value of a plan that has been tested before the emergency. A checklist can look solid on paper and still fail when people face pressure, conflicting signals and a fast-moving crisis. Companies are beginning to ask the same question about AI: what happens when an agent has to make decisions during a genuinely bad week?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get emergency and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s live experiment puts that question on public display. Its simulated company has synthetic employees and real money mechanics, and its AI participants face the same customers, crises and temptations. The point is to watch what they do, not just what they say they would do.

One company, five very different finishes

In the final Crucible League, published in July 2026, five models ran the same small software company through its worst week. The results ranged from 95 for gpt-5.6-sol to 73 for Opus 4.8, with Kimi K3 at 93, Sonnet 5 at 88 and Fable 5 at 77. A do-nothing baseline scored 26. The league’s rule is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

All the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding is a useful reminder for anyone planning for a breakdown: recognizing the emergency is not the same as carrying out the response. The experiment summed up the gap as “Same diagnosis, same pitch — no signature.”

The clue was already in the company’s files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It did not appear in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. In other words, the advantage came from finding and using relevant information that was already available, then following through.

That distinction matters in a business crisis. A team can spot the obvious warning and still miss the detail that changes the outcome. Firmulate’s experiment makes the decision trail watchable, with every decision versioned and auditable. Readers can also try to identify which model made real management decisions in a quiz built from 242 real, unedited decisions at Firmulate.

Pressure tests include trust and discipline

The manipulation test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” For organizations weighing AI access to customers, support queues or forecasts, that kind of refusal is part of preparedness: a system under pressure must preserve boundaries as well as identify threats.

Refusal alone, however, did not guarantee a strong result. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models. The lesson is not that a detailed analysis guarantees a good outcome. It is that a playbook also has to make the next safe and useful action clear.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also reports that the live company has 13 synthetic employees, burns €105k a month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Those mechanics make the experiment concrete, while keeping its company simulated.

From watching to a company’s own drill

For enterprises, Firmulate offers a pilot that applies the wargame to a read-only export of their own business. The export can provide a basis for testing crisis scenarios against company-specific information and examining weak points in existing playbooks. The exercise produces a board report with model rankings and findings about those weak points. Nothing writes back to real systems.

That boundary is central to the idea: rehearse decisions against a copy of the business before an AI agent is trusted with live operations. The public experiment shows the kind of gap a drill can reveal—spotting a crisis, refusing a trap, finding evidence and still failing to close the loop. The pilot lets a company ask how its own models and procedures handle that sequence.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Run the drill before the stakes are real

Preparedness depends on more than having a plan. It depends on knowing whether people and systems can find the right information, resist manipulation and follow a safe path to action when pressure rises. Firmulate’s live experiment makes those behaviors visible; its enterprise pilot brings the exercise to a company’s own business data.

To discuss a pilot using a read-only export, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The kids with phones are alright

Recent research shows children with phones are generally doing well, challenging concerns about digital device use and mental health.

The Role of Sensors in Next-Gen Survival Bots

Forces of innovation drive sensors in next-gen survival bots, unlocking unprecedented environmental awareness that could redefine their capabilities—discover how they adapt and survive.

Why Military Robots Are Shaping the Future of Defense

AIThis post was created with the assistance of artificial intelligence (AI).Military robots…

Utilizing Drones for Damage Assessment and Relief Efforts

Bringing advanced drone technology to damage assessment and relief efforts revolutionizes emergency response, but exploring further reveals how to maximize their potential.