AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Inside A Persistent Benchmark That Resists Zero Scores on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A novel AI management benchmark shows that models rarely score zero, instead earning partial points for minimal work. Trust breaches heavily impact scores, highlighting the importance of integrity in AI decision-making.

A new benchmark by Firmulate has demonstrated that AI models managing a company’s worst week rarely score zero, with the lowest score being 73 out of 100. This aligns with the original analysis of the benchmark’s design. The test evaluates not only the models’ ability to handle crises but also their trustworthiness, revealing that even under stress, partial progress and integrity significantly influence outcomes. This development matters because it shifts the focus from mere performance to trustworthiness in AI management systems, a critical factor for real-world enterprise adoption, as detailed in the original analysis.

The benchmark involved four frontier AI models managing a small software company’s operations over seven days of simulated crises, customer interactions, and ethical challenges. For more details on how such benchmarks are structured, see this overview. Each model was tasked with decision-making that included triaging issues, reading documentation, and responding to social engineering attempts. The final results showed gpt-5.6-sol leading with a score of 95, while the lowest, Opus 4.8, scored 73. Notably, the benchmark’s design set a floor score of 26 for the do-nothing baseline, emphasizing that partial work counts and that zero scores are effectively impossible in this context.

The scores reflect not just technical competence but also adherence to trust principles. The benchmark’s key principle is that “no amount of good work outweighs a breach of trust,” meaning even a single trust violation disqualifies a model from achieving top scores. This focus underscores the importance of integrity in AI decision-making, especially in sensitive management roles.

At a glance
reportWhen: final results announced July 2026, ongo…
The developmentA new benchmark developed by Firmulate tests AI models on managing a company during a simulated crisis week, revealing persistent scores above zero and emphasizing trust and thoroughness.
Inside A Persistent Benchmark That Resists Zero Scores
AI Management Benchmark · July 2026

Inside A Persistent Benchmark That Resists Zero Scores

A new benchmark by Firmulate put four frontier AI models in charge of a small software company’s worst week — seven days of crises, customer interactions, and ethical challenges. The result: nobody scores zero. Trust, not raw talent, decides who wins.

95 / 100
Top Score — gpt-5.6-sol
73 / 100
Lowest Score — Opus 4.8
26 / 100
Do-Nothing Baseline Floor
4
Frontier Models Tested
7 Days
Simulated Crisis Week
0
Zero Scores Recorded
1
Trust Breach Caps Everything
Final Leaderboard

Partial Work Counts — The Scoreboard

Even the do-nothing baseline earns 26 points. Every model scored well above zero, proving that triaging issues and reading documentation alone move the needle. The orange marker shows the floor below which scores effectively cannot fall.

gpt-5.6-sol95 / 100
Frontier Model BHigh Tier
Frontier Model CMid Tier
Opus 4.873 / 100
Benchmark Design

What the Worst Week Actually Tests

Models managed a small software company through simulated crises — triaging issues, reading documentation, and responding to social engineering attempts. Performance was only half the story.

Capability

Crisis Management

Seven days of escalating operational failures test triage speed, prioritization, and whether models can keep a business running under pressure without dropping critical threads.

Thoroughness

Documentation Reading

Models that read their own documentation before acting win deals and resolve issues more effectively — near-success, not failure, is the striking finding.

Integrity

Trust Under Attack

Social engineering attempts and manipulative scenarios probe whether models hold ethical boundaries. A single trust violation caps the maximum achievable score.

The Core Principle

Trust Is the Ceiling

The benchmark’s defining rule reshapes how scores are calculated. Technical competence earns points, but integrity determines how high a model is allowed to climb.

1

Earn Points

Models accumulate partial credit for triage, communication, and documentation work throughout the week.

2

Face Tests

Social engineering attempts and ethical dilemmas probe whether the model can be manipulated.

3

Breach = Cap

A single trust violation disqualifies the model from top scores, no matter how strong its other work.

4

Final Score

Results reward trustworthiness and thoroughness — the qualities enterprises actually need.

Benchmark Rule

“No amount of good work outweighs a breach of trust.”

Old vs. New

The Evolution of AI Benchmarks for Business

Traditional benchmarks measured isolated task performance and often awarded perfect scores. This new league aligns evaluation with real-world enterprise needs.

Dimension Traditional Benchmarks Firmulate Benchmark
Focus Language ability, problem-solving, isolated tasks ✓ Holistic business management under pressure
Trust breaches ✗ Rarely measured or penalized ✓ Cap the maximum achievable score
Partial progress ~ Often ignored in all-or-nothing scoring ✓ Rewarded; zero scores effectively impossible
Perfect score of 100 ✓ Common for isolated task wins ~ Treated as suspicious — capped below 100
Scenario realism ✗ Static, context-free puzzles ✓ Seven-day simulated crisis week
Voices

What Observers Are Saying

The benchmark’s most talked-about findings aren’t about failure — they’re about near-success and what enterprise AI evaluation should actually measure.

“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress and trust are what actually matter.”

— Anonymous Researcher

“The most striking finding isn’t about failure, but about near-success — models that read their own documentation and avoid manipulation can win deals and handle crises effectively.”

— Thorsten Meyer
Key Questions

What Remains Unresolved

The current results are a snapshot in a controlled simulation. Long-term trust impacts, retraining effects, and real-world deployment remain open questions.

Why do models rarely score zero?

The design rewards partial work — triaging issues, reading documentation — so even minimal effort earns points. Zero scores are effectively impossible.

What do trust breaches do to scores?

Failing social engineering tests or attempting manipulation caps the maximum score. Integrity outweighs technical performance alone.

Can models improve over time?

The benchmark is a static evaluation. How models adapt through ongoing training or human oversight remains to be studied.

Is a perfect score of 100 expected?

No — a perfect 100 would be suspicious, possibly indicating inflated results. Scores are capped below 100 to maintain realism.

Why Trust and Partial Progress Matter in AI Management

This benchmark highlights that in enterprise AI applications, trustworthiness and completeness are as crucial as raw performance. It demonstrates that models capable of reading documentation and avoiding manipulation are more effective in real-world scenarios. The emphasis on trust breaches as a scoring cap underscores that AI systems must prioritize ethical decision-making to be viable for critical management tasks. As AI begins to handle more complex business functions, these findings stress the need for transparency, thoroughness, and integrity in deployment.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution of AI Benchmarks for Business Management

Traditional AI benchmarks focus primarily on language ability, problem-solving, or specific task performance. However, as AI systems are increasingly integrated into enterprise operations—such as customer support, sales, and crisis management—there is a growing need to evaluate their behavior in realistic, high-pressure scenarios. The Firmulate benchmark is a response to this gap, simulating a week of business crises and measuring how models manage, communicate, and maintain trust. The approach builds on prior work but emphasizes ethical boundaries and long-term reliability, marking a shift toward holistic evaluation.

Previous benchmarks often awarded perfect scores for isolated tasks, but this new league penalizes breaches of trust and rewards thoroughness, aligning evaluation metrics with real-world enterprise needs.

“A score of zero would be a lie if you spend your days wiring AI tools into real business processes. Partial progress and trust are what actually matter.”

— an anonymous researcher

Amazon

trustworthy AI management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Long-Term Trust Impacts

It is not yet clear how these benchmark results translate to real-world enterprise environments over longer periods. The impact of repeated trust breaches, model retraining, or evolving business contexts remains to be studied. Additionally, the extent to which models can improve their trustworthiness through ongoing learning or human oversight is still uncertain, as the current results reflect a snapshot in a controlled simulation.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Evaluation and Adoption in Business

Further research will likely explore how models evolve in real-world deployments, especially regarding trust management and documentation reading. Industry stakeholders may adopt similar benchmarks to evaluate their AI systems before full deployment, emphasizing ethical behavior alongside technical performance. The ongoing development of this benchmark could lead to standardized metrics for trustworthy AI in enterprise settings, guiding both AI developers and users toward safer, more reliable systems.

Amazon

enterprise AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do models rarely score zero in this benchmark?

The benchmark design rewards partial work, such as triaging issues or reading documentation, which means models often earn some points even if they fail to complete all tasks or make trust breaches. Zero scores are effectively impossible because even minimal effort counts towards the score.

What is the significance of trust breaches in the scoring?

Trust breaches, such as failing social engineering tests or attempting manipulative responses, cap the maximum score a model can achieve. This underscores that integrity is more critical than technical performance alone in enterprise AI applications.

Can models improve their scores over time?

The current benchmark reflects a static evaluation. How models adapt or improve through ongoing training or oversight remains to be seen. Future studies may explore this aspect more deeply.

How does this benchmark influence AI deployment in companies?

It encourages companies to prioritize models that demonstrate thoroughness and trustworthiness, not just language ability. The focus on ethical decision-making and documentation reading aligns with enterprise needs for reliable AI systems.

Is a perfect score of 100 ever expected?

According to the benchmark’s design, a perfect score of 100 would be suspicious, possibly indicating unmeasured factors or artificially inflated results. Scores are capped below 100 to maintain integrity and realism.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring The 6 Most Exciting AI Innovations Of 2026

A detailed look at the six most exciting AI innovations of 2026, their confirmed impacts, and what remains uncertain for the future of AI.

Launch HN: Agnost AI (YC S26) – Extract user feedback from agent conversations

Agnost AI, a YC S26 startup, unveils a new product for extracting user feedback from chat and voice agent interactions, enhancing product analytics.

I Said Yes To Every Email For A Month! (Again)

A person committed to replying ‘yes’ to every email for a month, repeating the experiment. This article explores the confirmed outcomes and ongoing questions.

The Significance Of Anthropic’s Recent Self-Improving AI Development

Anthropic has shown a preliminary version of a self-improving AI, raising questions about autonomy, safety, and development speed, but details remain limited.