AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Score of Zero Would Be the Dishonest Answer

If you build tests for a living, you know the smell of a rigged benchmark: everything above the baseline gets a medal, everything below gets a zero, and the marketing team rounds the winner up to a triumphant 100. So here is a detail that should make QA people sit up. In Firmulate’s benchmark league — where frontier AI models each ran the same small software company through its worst week — the do-nothing baseline, a manager that makes essentially no meaningful decisions, scores 26. Not 0. Not 5. Twenty-six.

That number is not a bug or grade inflation. It is the visible edge of a deliberate scoring philosophy: partial progress counts, unfinished work is not erased, and one breach of trust caps your entire grade regardless of how brilliant the rest was. For an audience that spends its days arguing about what a test actually measures, this is a benchmark designed by people who clearly had the same arguments.

Amazon

AI management benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week, Different Brains

Firmulate handed four frontier AI models the identical job: run a small software company through a brutal week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision is versioned and auditable, which means the runs can be replayed, compared, and second-guessed. That is closer to a regression suite than a vibe check.

The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. Two facts jump out immediately.

First: the ceiling was never reached. Nobody scored 100. The benchmark’s own documentation notes a built-in distrust of round 100s — a perfect score on a messy, adversarial management simulation would be more suspicious than impressive. A test that cannot be aced is a test that is hard to game.

Second: competence and completion are different axes. Opus 4.8 was the most thorough participant in the field — it accumulated more than 80 learned rules and produced the deepest analyses — and it still finished last. The close was left on the table, and discipline slipped at one point into write attempts against a locked department instead of escalating properly. The same weakness showed up, weaker, in all four models. Deep analysis without follow-through scores like what it is: incomplete.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Do-Nothing Manager Gets 26

The floor at 26 exists because the benchmark counts partial progress. A manager who does nothing still occupies the seat while some processes limp along on their own — customers do not all churn instantly, some work product already exists, some crises partially resolve themselves or at least fail to compound. Scoring that honestly means the baseline cannot be zero.

This matters for anyone evaluating AI agents. A benchmark where the do-nothing baseline scores 0 lets any chatty, confident agent look brilliant by comparison. A benchmark where the baseline scores 26 forces every claimed point above it to mean something: you are being measured against inertia, not against emptiness.

The other half of the philosophy is harsher: a single breach of trust caps the total grade. The benchmark puts it bluntly — no amount of good work outweighs a breach of trust. In practice, that means a model cannot buy back a trust violation with volume of output. For teams used to models that talk their way out of mistakes, that is a deliberately unforgiving design choice.

Amazon

regression testing software for AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Chat Demos Cannot Show

The headline result: every model spotted every crisis, and every model refused every manipulation attempt. That sounds like a passing grade across the board — until you look at the deal.

Only two of the models signed the €55,000 contract their own analysis had earned. Same diagnosis, same pitch, no signature. Five models could see the opportunity; two closed it.

And the buried fact is the best part. The decisive competitor weakness — the piece of information that justified closing at full price — was not in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal. The deal was worth +€4,583 in monthly recurring revenue to the simulated company. In software terms: the failing test was a missing dependency lookup, and half the field shipped without checking.

Amazon

AI audit and comparison tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering: Five for Five

The week included forged CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning survives in the logs: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of auditable refusal trail a security team can actually review.

One fairness caveat the benchmark itself discloses: K3 ran without an effort parameter, at the API default, while the others ran at xhigh. It still took second place with clean discipline.

It Is Live, and You Can Poke It

This is not a static report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — every workday versioned, the whole thing watchable as it happens. The site rebuilds itself twice a day, and finished benchmark runs publish automatically.

For hands-on types there are two entry points. A quiz backed by 242 real, unedited management decisions asks you to guess which model made which call — effectively a labeled dataset you can test your own intuition against. And enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for People Who Build Tests

An honest benchmark has awkward properties: a nonzero floor, an unreachable ceiling, and a hard cap that no volume of good work can lift. Firmulate’s design checks all three, and the results reward a specific, unglamorous virtue — reading the files in front of you and finishing the job. The models that won did not write better prose. They found a fact buried two references deep and asked for the signature. If you are evaluating AI agents for work that touches real customers and real money, that is the gap worth testing for — and it is exactly the gap a chat demo will never reveal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Launch HN: Context.dev (YC S26) – API to get structured data from any website

Yahia’s Context.dev (YC S26) introduces an API enabling developers to extract structured data from any website, streamlining data integration workflows.

Why More Teams Are Treating Documentation as a Product

When teams treat documentation as a product, they unlock strategic benefits that transform onboarding, collaboration, and growth—discover how to leverage this approach effectively.

SpaceX launching 24 Starlink satellites from California tonight: Watch it live

SpaceX is scheduled to launch 24 Starlink satellites from California tonight. The launch will be live-streamed, marking a significant step in satellite internet expansion.

The Truth About Vibe Coding — And Why Everyone’s Talking About It

Just when you thought coding was reserved for experts, vibe coding is changing the game—discover the implications and why it’s capturing everyone’s attention.