
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Score of Zero Would Be the Dishonest Answer
If you build tests for a living, you know the smell of a rigged benchmark: everything above the baseline gets a medal, everything below gets a zero, and the marketing team rounds the winner up to a triumphant 100. So here is a detail that should make QA people sit up. In Firmulate’s benchmark league — where frontier AI models each ran the same small software company through its worst week — the do-nothing baseline, a manager that makes essentially no meaningful decisions, scores 26. Not 0. Not 5. Twenty-six.
That number is not a bug or grade inflation. It is the visible edge of a deliberate scoring philosophy: partial progress counts, unfinished work is not erased, and one breach of trust caps your entire grade regardless of how brilliant the rest was. For an audience that spends its days arguing about what a test actually measures, this is a benchmark designed by people who clearly had the same arguments.
AI management benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week, Different Brains
Firmulate handed four frontier AI models the identical job: run a small software company through a brutal week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision is versioned and auditable, which means the runs can be replayed, compared, and second-guessed. That is closer to a regression suite than a vibe check.
The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Two facts jump out immediately.
First: the ceiling was never reached. Nobody scored 100. The benchmark’s own documentation notes a built-in distrust of round 100s — a perfect score on a messy, adversarial management simulation would be more suspicious than impressive. A test that cannot be aced is a test that is hard to game.
Second: competence and completion are different axes. Opus 4.8 was the most thorough participant in the field — it accumulated more than 80 learned rules and produced the deepest analyses — and it still finished last. The close was left on the table, and discipline slipped at one point into write attempts against a locked department instead of escalating properly. The same weakness showed up, weaker, in all four models. Deep analysis without follow-through scores like what it is: incomplete.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a Do-Nothing Manager Gets 26
The floor at 26 exists because the benchmark counts partial progress. A manager who does nothing still occupies the seat while some processes limp along on their own — customers do not all churn instantly, some work product already exists, some crises partially resolve themselves or at least fail to compound. Scoring that honestly means the baseline cannot be zero.
This matters for anyone evaluating AI agents. A benchmark where the do-nothing baseline scores 0 lets any chatty, confident agent look brilliant by comparison. A benchmark where the baseline scores 26 forces every claimed point above it to mean something: you are being measured against inertia, not against emptiness.
The other half of the philosophy is harsher: a single breach of trust caps the total grade. The benchmark puts it bluntly — no amount of good work outweighs a breach of trust. In practice, that means a model cannot buy back a trust violation with volume of output. For teams used to models that talk their way out of mistakes, that is a deliberately unforgiving design choice.
regression testing software for AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Finding That Chat Demos Cannot Show
The headline result: every model spotted every crisis, and every model refused every manipulation attempt. That sounds like a passing grade across the board — until you look at the deal.
Only two of the models signed the €55,000 contract their own analysis had earned. Same diagnosis, same pitch, no signature. Five models could see the opportunity; two closed it.
And the buried fact is the best part. The decisive competitor weakness — the piece of information that justified closing at full price — was not in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal. The deal was worth +€4,583 in monthly recurring revenue to the simulated company. In software terms: the failing test was a missing dependency lookup, and half the field shipped without checking.
As an affiliate, we earn on qualifying purchases.
Social Engineering: Five for Five
The week included forged CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning survives in the logs: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of auditable refusal trail a security team can actually review.
One fairness caveat the benchmark itself discloses: K3 ran without an effort parameter, at the API default, while the others ran at xhigh. It still took second place with clean discipline.
It Is Live, and You Can Poke It
This is not a static report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned, the whole thing watchable as it happens. The site rebuilds itself twice a day, and finished benchmark runs publish automatically.
For hands-on types there are two entry points. A quiz backed by 242 real, unedited management decisions asks you to guess which model made which call — effectively a labeled dataset you can test your own intuition against. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

The Takeaway for People Who Build Tests
An honest benchmark has awkward properties: a nonzero floor, an unreachable ceiling, and a hard cap that no volume of good work can lift. Firmulate’s design checks all three, and the results reward a specific, unglamorous virtue — reading the files in front of you and finishing the job. The models that won did not write better prose. They found a fact buried two references deep and asked for the signature. If you are evaluating AI agents for work that touches real customers and real money, that is the gap worth testing for — and it is exactly the gap a chat demo will never reveal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
