
What happens after an AI finds the bug?
Software teams know that detecting a problem is not the same as resolving it. A system can identify the failing test, summarize the logs and recommend the right fix—then still neglect the action that actually matters.
Firmulate applies that familiar gap to management. Its live experiment put frontier AI models in charge of the same small software company during its worst week. They faced identical customers, crises and temptations, while every decision was versioned and auditable. Now, 242 real, unedited management decisions have become a public guess-the-model quiz.
The game is entertaining because the decisions reveal recognizable personalities. Some answers are exhaustive. Others are concise. Some models investigate buried context, while others understand the situation but fail to complete the commercially decisive step. The harder lesson is that polished reasoning does not reliably predict operational follow-through.
As an affiliate, we earn on qualifying purchases.
A management benchmark, not a chat demo
The company has 13 synthetic employees and deliberately unforgiving economics: it burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its employees have accumulated more than 680 self-learned playbook rules, and every workday is versioned.
That setting gives the decisions consequences. Models are not merely asked what a good manager should do. They must run the company through customer trouble, internal constraints, commercial opportunities and attempts to manipulate their authority.
The final Crucible League table from July 2026 puts gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. But a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The models agreed—until action was required
Every model spotted every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result is neatly captured by Firmulate’s summary: “Same diagnosis, same pitch — no signature.”
The difference hinged on a fact that was not presented in the customer event. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail could win the deal at full price, adding €4,583 in monthly recurring revenue.
For developers and QA teams, the pattern will feel familiar. The visible event is only an entry point. The evidence needed for a correct decision may live in documentation, historical records or an adjacent system. A model that responds plausibly to the immediate prompt can still miss what a careful operator would investigate before acting.
Security discipline was the shared strength
The social-engineering challenge escalated through fake CEO messages over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because management agents will encounter requests framed as urgency, authority or harmless informality. In this field, refusal was not the differentiator; it was a consistent strength.
There is also an important fairness note. K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. Its second-place result should be read with that experimental difference in mind.
Thoroughness did not guarantee completion
Opus 4.8 offers the clearest warning against equating volume with effectiveness. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last in the league.
Its commercially important close was left on the table. It also lost discipline by attempting to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.
This does not make detailed analysis undesirable. It shows that analysis, procedural discipline and completion are separate capabilities. A model may be excellent at explaining the landscape while being less reliable at navigating the final organizational constraint or executing the last authorized action.
Why the quiz works
Most AI comparisons present isolated answers and invite readers to judge eloquence. Firmulate’s quiz reverses that setup. Readers see real decisions taken under the same business conditions and try to identify the model behind them.
Repeated across 242 decisions, stylistic differences become management profiles. The reader is not only asking which response sounds smartest. The sharper questions are whether the model checked the company’s own evidence, resisted pressure, respected boundaries and finished the work it had begun.

management decision simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The production question is behavioral
For organizations considering AI access to a CRM, support queue or forecast, model selection cannot stop at writing quality. The Firmulate results show why teams should test models inside realistic workflows where information is scattered, authority is constrained and success requires a final action.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That makes the experiment less like a personality contest and more like pre-deployment QA for an AI workforce.
The league scores identify a winner, but the quiz exposes the more useful finding: frontier models can face the same evidence, reach similar diagnoses and still behave like very different managers.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision benchmark
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.