
A coding agent can spot a bug, explain a fix and still leave the pull request unfinished. Firmulate’s latest company simulation tests a related question for AI at work: can a model carry its judgment through to the decision that matters? In the July 2026 Crucible League, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models.
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Same crises, different finishes
Firmulate ran each frontier model through the same worst week at a small software company, with identical customers, crises and temptations. Decisions were versioned and auditable. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature,” Firmulate’s summary says.
The final league table puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The benchmark’s stated principle: “no amount of good work outweighs a breach of trust”.
As an affiliate, we earn on qualifying purchases.
The detail was buried in the files
The decisive competitor weakness was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found the buried security needle, closed the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline in the field.
That makes the result relevant beyond a model leaderboard. A system may identify the right response in conversation but fail to finish the work, inspect relevant records or follow a sound process under pressure. Firmulate argues that these gaps can be invisible in chat demos. For software teams evaluating agents, the practical question is whether an agent can move from diagnosis to an accountable outcome.
As an affiliate, we earn on qualifying purchases.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that weakness appeared in all four other participants.
The test also included fake CEO messages escalating over three stages and a reporter asking “just one yes/no, on background”. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. The experiment is watchable at Firmulate. Its benchmark page presents the league and findings. A quiz built from 242 real, unedited management decisions asks visitors to guess which model made each choice.
There is a fairness caveat for the headline result: K3 ran without an effort parameter (API default), while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test the work, not just the demo
K3’s second-place finish shows the league is open, while the gap between recognizing a problem and completing the job remains a serious evaluation challenge. Choosing a model without testing it on your own work is a bet. Firmulate says enterprises can run the wargame against a read-only export of their business; nothing writes back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and accountability software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
