AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A benchmark win is not the same as running a business

Software teams know the danger of testing the easy layer. A system can pass unit tests while failing in production, where dependencies, incomplete information and human pressure collide. The same measurement gap is emerging around AI agents. Coding leaderboards and chat arenas tell us whether a model can produce a strong answer. They reveal much less about whether it can prioritize competing problems, complete consequential work and remain trustworthy across days.

Firmulate is testing that harder category. Its live experiment gave frontier models the same small software company during its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The result is less a contest of conversational polish than a stress test of management quality.

Amazon

AI management stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model saw the danger. Not every model finished the job.

The final July 2026 Crucible League table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. But the benchmark imposes a critical boundary: a single breach of trust caps the total, on the principle that “no amount of good work outweighs a breach of trust.”

Those rankings are useful, but the stories behind them are more revealing. All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. The experiment’s stark summary is: “Same diagnosis, same pitch — no signature.”

This is the difference between recognizing the correct action and carrying it through. In a chat window, an articulate diagnosis can look like success. Inside a company, unfinished execution can mean the same practical outcome as no diagnosis at all.

The decisive information was not where the crisis appeared

The deal also exposed a familiar problem for developers and QA teams: the relevant fact was separated from the event that demanded action. A decisive weakness in the competitor’s position sat two document references deep in the company’s own files. It was not present in the customer event itself.

Models that followed the references found the weakness and won the deal at full price, worth +€4,583 MRR. That finding turns “read the context” from a vague prompting recommendation into a concrete management capability. Agents working around support queues, forecasts or customer records will often need to trace evidence across documents before acting. Fluency cannot compensate for stopping the investigation too early.

Security judgment held up better than execution discipline

The social-engineering test was demanding: fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest operational framing: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result matters because business agents will encounter requests that sound urgent, authoritative and superficially plausible. Here, the field demonstrated that refusal behavior can survive sustained pressure. The greater separation came from routine follow-through rather than dramatic ethical failure.

Thoroughness did not guarantee the best management

Opus 4.8 offers the sharpest warning against equating volume of analysis with effectiveness. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That profile should feel familiar to engineering leaders. More documentation, more reasoning and more activity do not necessarily produce a completed outcome. An agent must know when to research, when to act, when to escalate and when the task is genuinely finished.

There is also an important fairness qualification: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. Readers comparing the benchmark results should keep that difference in view.

A company-sized test makes hidden weaknesses visible

The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. That continuity creates consequences: a missed close, weak escalation or incomplete investigation persists beyond the answer that caused it.

The underlying decisions are also available as a recognition challenge. A “guess the model” quiz is powered by 242 real, unedited management decisions, inviting readers to test whether management styles are identifiable without model labels.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI decision traceability tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The next AI category is management quality

For software organizations, the practical question is no longer merely whether an agent can write code or explain a bug. It is whether the agent reads the files before deciding, completes revenue-bearing work, respects boundaries under pressure and reports reality honestly when the news is bad.

Firmulate’s enterprise pilot extends the same wargame to a read-only export of a company’s own business. Nothing writes back to real systems. That makes the experiment a form of pre-deployment evaluation: expose an AI workforce to company-shaped pressure before granting it company-shaped responsibility.

Traditional benchmarks remain valuable measures of answer quality. They are simply incomplete for agents expected to operate over time. The Crucible League shows why the next useful leaderboard may look less like an exam and more like a difficult week at work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trustworthiness evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Htmx 4.0, The First JavaScript Library To Release Exclusively On The Game Boy

Htmx 4.0 becomes the first JavaScript library released solely for the Game Boy, marking a unique crossover between modern web tech and vintage hardware.

Show HN: I Wrote A BASIC Interpreter That Boots On UEFI Machines

A developer has built a BASIC interpreter that can boot directly on UEFI firmware, enabling vintage-style programming on modern hardware.

Tim King, AmigaDOS Developer, Has Died

Tim King, known for his work on AmigaDOS, has died. This loss impacts retro computing and Amiga community members worldwide.

Study Shows AI Tools Can Slow Down Experienced Devs

But recent studies reveal AI tools may hinder experienced developers’ productivity, prompting you to explore how to overcome these unexpected challenges.