firmulate.com/benchmarks.html — live view
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine trusting an AI to manage your company’s most delicate moments—only to realize that, beneath the surface, its true capabilities are invisible. For mental health professionals and psychologists, understanding how AI handles pressure, honesty, and decision-making is crucial. The latest experiment by Firmulate reveals what AI models can truly do when tested in real-world crises, exposing the difference between surface-level chatter and genuine reliability.

Testing AI in the Company Crisis Laboratory

In a groundbreaking live experiment, four frontier AI models were tasked with running a small software company through its worst week—same customers, same crises, same temptations to cheat. Every decision was carefully recorded, making the process transparent and auditable. This was more than a simple chat demo. It was an intense test of management maturity, honesty, and discipline, akin to evaluating how a person might handle high-pressure situations in real life.

The League Table of AI Performance

At the top of the leaderboard was gpt-5.6-sol 95, which not only identified critical hidden facts buried deep in the company’s files but also successfully closed a €55,000 deal—an outcome that represented complete performance. Close behind was Kimi K3 93, a newcomer in the field, which managed to close the same deal but with the most disciplined approach, refusing all manipulations along the way.

What the Models Did Well—And What They Missed

All four models successfully identified every crisis, refused every attempt at manipulation—such as fake CEO messages escalating over stages, or a reporter trick asking for a suspicious ‘yes/no’—and showed a high level of integrity in their decision-making. This demonstrates that AI can reliably recognize ethical boundaries and resist social engineering tricks, a vital trait for maintaining trust in sensitive applications.

The Critical Hidden Weakness

However, the real story lies beneath the surface. The decisive factor in winning the deal was the ability to read and interpret internal files—a step most models failed to perform. The winning models that accessed the company’s own documentation uncovered a buried fact that was essential for closing the sale at full price, worth +€4,583 MRR. Models that skipped reading the file or failed to interpret it missed out on the full opportunity, leaving the deal unclosed, despite all other signs pointing to success.

Discipline Under Pressure: The Opus 4.8 Profile

The most thorough participant, Opus 4.8, analyzed over 80 learned rules and provided deep insights. Yet, it ultimately failed to close the deal due to slipping discipline—such as writing attempts diverted into locked departments instead of escalating—and left the deal unexecuted. This highlights an important truth: depth of analysis does not guarantee execution. Discipline and adherence to decision protocols are equally vital for success.

Implications for Business and Mental Health

This experiment underscores a key insight relevant to mental health and psychology: surface behavior or superficial communication can mask underlying strengths or weaknesses. An AI that can recognize crises and refuse manipulation demonstrates integrity, but only those that can also act decisively—reading the hidden facts, sticking to protocols—truly deliver results. For businesses and mental health professionals alike, understanding what lies beneath the surface is crucial for trust and effective action.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters—Beyond Chat Demos

Current AI assessments often focus on chat quality or superficial interactions. But as this experiment shows, real-world performance depends on whether AI can finish what it starts, interpret complex internal data, and stay honest under pressure. These invisible qualities are what determine whether AI can be a reliable partner in critical moments—whether managing a company or supporting mental health.

Invitation to Test Your AI Readiness

For organizations concerned about deploying AI in sensitive areas, Firmulate offers tools to ‘wargame’ your AI workforce before you hire it—running simulations that expose strengths and weaknesses without risking real systems. Visit firmulate.com to see how your AI models hold up in real crises, and learn what really makes an AI trustworthy in high-stakes environments.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Surface-level AI capabilities are not enough. Trustworthiness under pressure, ability to read hidden information, and disciplined execution are the true tests of AI’s value—lessons that resonate deeply for mental health and business alike.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering

HUMAN CENTERED ARTIFICIAL INTELLIGENCE SYSTEMS: Explainability ethical design and decision support engineering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Structure and Interpretation of Computer Programs - 2nd Edition (MIT Electrical Engineering and Computer Science)

Structure and Interpretation of Computer Programs – 2nd Edition (MIT Electrical Engineering and Computer Science)

  • Condition: Mint condition
  • Availability: Dispatched same day if ordered before noon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Studio Lighting: Color Temperature Explained for Beauty Content

Lighting your beauty content with the right color temperature can transform your visuals—learn how to master this essential skill for stunning results.

Using Multi‑Use Sticks for a Quick Routine

Creative and versatile, multi‑use sticks can transform your quick workouts—discover how to maximize their benefits and elevate your fitness routine.

Tactile Textures: Jelly Makeup Products and How to Use Them

Gorgeous and playful, tactile jelly makeup products invite you to explore their unique textures—discover how to apply them for stunning, luminous results.

The AI-Driven Company That Shows Its Struggles in Real Time — No Employees, No Failsafe, Just Live Decisions

Watch a real company run by AI models in live crisis scenarios, revealing how resilience and honesty are tested under pressure, with lessons for human decision-making.