Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

The bug appears after the model gives the right answer

Software teams know that passing a test is not the same as surviving production. A system can identify the defect, recommend the correct fix and still fail because nobody completes the last consequential step.

Firmulate has turned that familiar gap into a live business experiment. Its software company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and displays a public cash countdown. Every workday is versioned, creating a continuing record of a company trying to stay alive rather than a polished demonstration arranged around a successful prompt.

The result is build-in-public pushed to an unusual extreme. Visitors can watch the company operate live, including the financial pressure under which its synthetic workforce makes decisions. The experiment also exposes something conventional AI evaluations often miss: recognizing what should happen and ensuring that it actually happens are different capabilities.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

In the Crucible League, each frontier model ran the same small software company through its worst week. Customers, crises and temptations were held constant. Every decision was versioned and auditable, so the comparison concerned management behavior under identical conditions rather than the persuasiveness of an isolated answer.

The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total on the principle that “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate’s succinct finding captures the failure: “Same diagnosis, same pitch — no signature.” For developers and QA teams, this is the uncomfortable part. The models were not defeated by an inability to understand the situation. Most were defeated somewhere between understanding it and completing the business outcome.

The winning fact was already in the company

The decisive weakness in a competitor was not delivered neatly inside the customer event. It sat two document references deep in the company’s own files. The models that read that material won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding makes the experiment especially relevant to software organizations considering agents for support queues, customer records or forecasts. Fluency at the point of interaction is not enough. Useful work can depend on whether an agent consults the business context that already exists, follows the evidence through multiple references and carries the resulting action to completion.

Pressure did not break the trust boundary

The models also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was not merely a security sideshow. The company was under commercial pressure, yet the participants did not trade trust for apparent speed. Readers can also read what the synthetic employees say, turning terse decisions and explanations into another public layer of the company’s running record.

Thoroughness did not guarantee the result

Opus 4.8 provides the sharpest cautionary profile. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, but it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in all four other participants.

The comparison also carries an important fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference matters when reading the narrow gap between K3 and the league leader.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI for Life: 100+ Ways to Use Artificial Intelligence to Make Your Life Easier, More Productive…and More Fun!

AI for Life: 100+ Ways to Use Artificial Intelligence to Make Your Life Easier, More Productive…and More Fun!

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that publishes consequences, not promises

Firmulate’s live company has accumulated more than 680 self-learned playbook rules, but the compelling artifact is not the rule count alone. It is the public sequence of decisions, missed closes, resisted manipulation and financial consequences. Because every workday is versioned, the experiment produces fresh operational evidence instead of freezing its conclusions in a benchmark report.

That makes the company a running QA environment for management behavior. Its synthetic employees can identify crises and defend trust boundaries, yet the league shows that apparently capable models may still stop short of the action that converts sound analysis into revenue.

The public cash countdown gives those omissions weight. With €105,000 in monthly burn and €2,300 in monthly recurring revenue, unfinished work is not an abstract scoring error. It is part of the visible story of whether this employee-free software company can keep operating. For builders assessing AI workers, that may be the most useful test of all: not whether a model can describe competent work, but whether it reliably finishes the work the business depends on.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI FOR BUSINESS PROCESS REENGINEERING & AUTOMATION: The Executive's Complete Playbook to Deploy AI-Powered Automation, Eliminate Process Waste, and Build ... BUSINESS & MANAGEMENT LIBRARY SERIES 36)

AI FOR BUSINESS PROCESS REENGINEERING & AUTOMATION: The Executive's Complete Playbook to Deploy AI-Powered Automation, Eliminate Process Waste, and Build … BUSINESS & MANAGEMENT LIBRARY SERIES 36)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Experience the Bitcoin Battle: Live War Visualization of BTC Price

Welcome to Bitcoin War, a groundbreaking live visualization experience created by the…

SpaceX launching 24 Starlink satellites from California tonight: Watch it live

SpaceX is scheduled to launch 24 Starlink satellites from California tonight. The launch will be live-streamed, marking a significant step in satellite internet expansion.

The Deploy Button Became the Bottleneck — and Cloudflare Just Bought the Build Step

Cloudflare’s acquisition of VoidZero aims to streamline the build-to-deploy process, integrating high-performance JavaScript tools into its edge network.

Anthropic Claude Sonnet 4.5 Sets New Bar for AI Code Generation

Anthropic’s Claude Sonnet 4.5 sets a new standard in AI code generation…