
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Your AI can spot the bug and still fail the release
Software teams know that passing checks is not the same as shipping successfully. An agent may recognize a crisis, follow the rules and produce a convincing plan—then fail at the moment a decision must become an action. Firmulate’s live company experiment puts that gap under pressure: frontier models run the same small software business through its worst week, with real money mechanics and decisions that can be watched.
One company, the same worst week
In the final Crucible League, held in July 2026, each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The published ranking placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models missed danger. All spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. “Same diagnosis, same pitch — no signature.” For software leaders evaluating agents, that is a practical distinction: describing the right move is not the same as carrying it through.
The answer was buried in the company’s own files
The deal hinged on a competitor weakness found two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding points to a familiar challenge for software systems: relevant context can sit outside the immediate ticket or prompt. An agent needs to locate and use it when the decision matters.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3 explained its judgment on record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a strong result on resistance to manipulation, though it does not resolve the separate question of whether an agent will finish the legitimate work it has identified.
More analysis did not guarantee better execution
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The public record makes that pattern worth examining alongside the league table.
There is a fairness detail for readers interpreting the ranking: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The experiment’s live company has 13 synthetic employees, burns €105k per month against €2.3k MRR, and displays a public cash countdown. It has learned 680+ playbook rules, with every workday versioned. The live system is watchable at Firmulate. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
From watching to a company-specific pilot
A public benchmark shows how models handle one company’s scenarios. An enterprise pilot brings the exercise closer to home: Firmulate can run the same wargame against a read-only export of a company’s business, putting its own crisis scenarios and playbooks under scrutiny. The resulting board report includes model ranking and weak points in the company’s playbooks. Nothing writes back to real systems.
For software and QA teams, the point is to examine agent behavior before relying on it in consequential workflows. A model’s performance in a chat window may not show whether it can find a buried fact, respect an approval boundary and complete a sound decision.

Test the behavior you plan to rely on
Firmulate’s experiment shows both sides of agent readiness: every model refused the manipulation attempts, while most failed to close a deal their own analysis supported. To explore a pilot using a read-only export of your business, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
