
In a crisis, recognizing what is wrong and taking the next useful step are different abilities. That gap matters in mental health, where a careful assessment can help only if it leads to timely support. It also matters when a company hands decisions to AI: a model may diagnose a business problem and still fail to follow through.
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A company’s worst week, repeated
Firmulate puts AI models in charge of a small software company facing a deliberately difficult week: the same customers, crises and temptations for each model. The live experiment is real and watchable. Its company has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the pressure visible.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Diagnosis is not the same as follow-through
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” In a chat exchange, the models’ explanations might sound equally convincing. The difference emerged when they had to carry a decision through.
The deal depended on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. It is a small but revealing distinction: the decisive clue was available, but finding it required attention to the company’s own records.
The experiment also tested pressure through fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear boundary under pressure; the unfinished deal shows that sound judgment in one moment does not guarantee action in another.
Thorough work can still leave the close open
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped: the model attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
That result makes the story more useful than a simple ranking. A model can show care, refuse manipulation and build a detailed account of events, yet stumble on execution. For organizations thinking about AI and people-facing work, that distinction deserves attention: reliability includes what happens after a risk is recognized.
Firmulate’s live company has accumulated more than 680 self-learned playbook rules, with every workday versioned. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. The public experiment offers a way to watch these decisions unfold rather than judge a model from polished answers alone.
From watching to a company-specific pilot
The experiment is a public demonstration; an enterprise pilot applies the wargame to a read-only export of its own business. The resulting board report can rank models and identify weak points in the company’s playbooks. Nothing writes back to real systems. That lets an organization examine how models handle its customers, pressures and decisions before considering a role for them in actual work.
For mental health and psychology readers, the lesson is not that AI can replace human judgment. It is that recognizing a problem is only one part of responsible action. Firmulate makes the gap observable in a company setting, where the decisions and their consequences can be reviewed.

See how AI handles pressure in the live experiment at Firmulate. To wargame your own company using a read-only export, explore the enterprise pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
