firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In a crisis, recognizing what is wrong and taking the next useful step are different abilities. That gap matters in mental health, where a careful assessment can help only if it leads to timely support. It also matters when a company hands decisions to AI: a model may diagnose a business problem and still fail to follow through.

For listenersOffer from Amazon

Turn your wind-down time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate puts AI models in charge of a small software company facing a deliberately difficult week: the same customers, crises and temptations for each model. The live experiment is real and watchable. Its company has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the pressure visible.

The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Diagnosis is not the same as follow-through

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” In a chat exchange, the models’ explanations might sound equally convincing. The difference emerged when they had to carry a decision through.

The deal depended on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. It is a small but revealing distinction: the decisive clue was available, but finding it required attention to the company’s own records.

The experiment also tested pressure through fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear boundary under pressure; the unfinished deal shows that sound judgment in one moment does not guarantee action in another.

Thorough work can still leave the close open

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped: the model attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

That result makes the story more useful than a simple ranking. A model can show care, refuse manipulation and build a detailed account of events, yet stumble on execution. For organizations thinking about AI and people-facing work, that distinction deserves attention: reliability includes what happens after a risk is recognized.

Firmulate’s live company has accumulated more than 680 self-learned playbook rules, with every workday versioned. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice. The public experiment offers a way to watch these decisions unfold rather than judge a model from polished answers alone.

From watching to a company-specific pilot

The experiment is a public demonstration; an enterprise pilot applies the wargame to a read-only export of its own business. The resulting board report can rank models and identify weak points in the company’s playbooks. Nothing writes back to real systems. That lets an organization examine how models handle its customers, pressures and decisions before considering a role for them in actual work.

For mental health and psychology readers, the lesson is not that AI can replace human judgment. It is that recognizing a problem is only one part of responsible action. Firmulate makes the gap observable in a company setting, where the decisions and their consequences can be reviewed.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how AI handles pressure in the live experiment at Firmulate. To wargame your own company using a read-only export, explore the enterprise pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Using Multi‑Use Sticks for a Quick Routine

Creative and versatile, multi‑use sticks can transform your quick workouts—discover how to maximize their benefits and elevate your fitness routine.

Berry Blush and Lips: Monochrome Looks

Aiming for the perfect monochrome berry look? Discover how seamless blending and shade choices can transform your beauty routine.

Soft Focus Skin: Combining Matte and Luminous Finishes

Aiming for flawless soft focus skin? Discover how blending matte and luminous finishes creates a radiant, natural look you’ll want to master.

Toning Down Bold Colors: Everyday Wearable Neon

Harness the power of neutral tones to effortlessly incorporate bold neon colors into your everyday wardrobe and discover the secrets to stylishly balancing vibrancy.