AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A coding agent can spot a bug, explain a fix and still leave the pull request unfinished. Firmulate’s latest company simulation tests a related question for AI at work: can a model carry its judgment through to the decision that matters? In the July 2026 Crucible League, Moonshot’s Kimi K3 placed second, ahead of three Western frontier models.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Same crises, different finishes

Firmulate ran each frontier model through the same worst week at a small software company, with identical customers, crises and temptations. Decisions were versioned and auditable. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. “Same diagnosis, same pitch — no signature,” Firmulate’s summary says.

The final league table puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The benchmark’s stated principle: “no amount of good work outweighs a breach of trust”.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The detail was buried in the files

The decisive competitor weakness was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found the buried security needle, closed the deal and saved the churning customer. It resisted all three baits and had one deviation, the cleanest discipline in the field.

That makes the result relevant beyond a model leaderboard. A system may identify the right response in conversation but fail to finish the work, inspect relevant records or follow a sound process under pressure. Firmulate argues that these gaps can be invisible in chat demos. For software teams evaluating agents, the practical question is whether an agent can move from diagnosis to an accountable outcome.

Amazon

enterprise AI agent tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that weakness appeared in all four other participants.

The test also included fake CEO messages escalating over three stages and a reporter asking “just one yes/no, on background”. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. The experiment is watchable at Firmulate. Its benchmark page presents the league and findings. A quiz built from 242 real, unedited management decisions asks visitors to guess which model made each choice.

There is a fairness caveat for the headline result: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI workflow automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not just the demo

K3’s second-place finish shows the league is open, while the gap between recognizing a problem and completing the job remains a serious evaluation challenge. Choosing a model without testing it on your own work is a bet. Firmulate says enterprises can run the wargame against a read-only export of their business; nothing writes back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and accountability software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Open-Source AI Models Gain Ground: LLaMA 3 and Beyond in 2025

Open-source AI models like LLaMA 3 are transforming AI development in 2025—discover how this shift could impact innovation and ethics.

Is Avoiding AI Distillation A Smart Move? ByteDance Weighs In

ByteDance’s Seed team commits to avoiding AI model distillation, even if it slows AI development, signaling a stance on training practices amid industry disputes.

AR & VR Development in 2025: Current State and Outlook

Harness the evolving AR and VR landscape of 2025 to discover how immersive experiences will redefine technology and daily life.

The Last MPEG-4 Visual Patent Has Expired

The final MPEG-4 Visual patent has officially expired, ending patent claims on the standard and potentially impacting licensing and technology use.