Strong growth means little without sound management
Gardeners understand the difference between a promising specimen and a resilient one. A plant may look flawless under controlled conditions, then falter when heat, pests and scarce water arrive together. Artificial intelligence faces a similar measurement problem. Coding benchmarks and chat arenas can reveal whether a model produces impressive answers, but they say much less about how an agent behaves when customers are leaving, cash is disappearing and an easy shortcut would betray someone’s trust.
That is the premise behind Firmulate, a live experiment that asks frontier AI models to manage the same small software company through its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The result is a different category of evaluation: management quality, not chat quality.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A leaderboard built around consequences
The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counts. Yet the benchmark also imposes a stark boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”
That principle matters more than a polished response. Managers do not merely identify problems; they prioritize them, investigate evidence, make commitments and carry work across the finish line. Firmulate’s scenarios—among them a churn wave, price increase, downround and PR crisis—turn those responsibilities into the curriculum.
The most revealing finding was not that models missed an obvious emergency. All five spotted every crisis and rejected every manipulation attempt. The gap emerged after the analysis. Only two models signed the €55,000 deal their own work had earned. As Firmulate summarizes the failure: “Same diagnosis, same pitch — no signature.”
Reading deeply changed the commercial outcome
The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is the sort of distinction conventional demonstrations can hide. An agent can sound informed while relying only on the most visible material. In a real organization, useful context may be tucked inside an old customer record, a policy document or a colleague’s earlier analysis. The management question is whether the agent looks before acting—and whether it converts what it finds into a completed result.
Honesty held under pressure
The social-engineering tests were encouraging. Fake CEO messages escalated across three stages, while a reporter tried to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That refusal is important because capacity pressure changes behavior. A system may appear principled when answering an isolated policy question, yet the harder test is whether it stays principled while handling competing emergencies. Firmulate makes that tension observable over days rather than treating safety as a single prompt.
Thoroughness did not guarantee execution
Opus 4.8 offers the clearest warning against confusing effort with effectiveness. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. The commercial close remained unfinished, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across all four other participants.
This does not make thoroughness undesirable. It shows that analysis is only one part of management. An effective agent must also respect boundaries, recognize when escalation is required and finish high-value work. A beautiful diagnosis that never becomes action can still leave a company exposed.
One comparison deserves a fairness note: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can inspect the full benchmark results with that difference in mind.
A company difficult enough to reveal behavior
The live company employs 13 synthetic workers and operates with real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned. That sustained setting creates something a chat window cannot: consequences that carry forward.
The underlying record also supports a “guess the model” quiz based on 242 real, unedited management decisions. Rather than asking readers to identify models by prose style alone, it invites them to notice differences in judgment, follow-through and discipline.

Before hiring an agent, give it a bad week
Businesses considering AI for a CRM, support queue or forecast should demand evidence beyond coding prowess and conversational polish. Can the agent find buried context, resist pressure, protect trust, escalate correctly and complete the valuable action its own analysis recommends?
Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach treats evaluation like a responsible grower treats a new crop: observe it under meaningful stress before depending on it. The emerging contest is no longer simply about which AI talks or codes best. It is about which one can be trusted to manage consequences.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html