🔍 Read the full analysis: The True Cost Of Diligent AI Failures on ThorstenMeyerAI.com
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
Start your free trialAs an affiliate, we earn on qualifying purchases.
TL;DR
An ongoing experiment demonstrates that AI models can thoroughly analyze business scenarios but still fail at the final step of execution. This exposes a critical gap in automation that could cost companies millions. The findings raise questions about how businesses should evaluate AI performance beyond intelligence and thoroughness.
A recent live experiment by Firmulate has demonstrated that even the most thorough AI models can fail at the critical final step of executing business decisions, despite excellent analysis and crisis recognition. This failure underscores a significant challenge for companies relying on AI automation: thoroughness does not guarantee operational success or business impact.
The experiment involved five AI models operating within a simulated company environment, where each was tasked with handling crises, analyzing opportunities, and closing deals. The most diligent model, Opus 4.8, identified all crises, resisted manipulative tactics, and produced comprehensive analyses, including 80 learned rules and detailed playbook insights. Despite this, Opus finished last in deal closure, with only 73 points out of a possible higher score, and failed to sign a major €55,000 contract. In contrast, two other models, including Kimi K3, successfully closed the deal, adding €4,583 in monthly revenue.
This discrepancy was traced to a subtle but decisive weakness: Opus failed to act on a critical piece of information buried two documents deep in the company’s files, which if used, could have supported the sale. The models that followed that trail successfully closed the deal, illustrating that the final, decisive step often hinges on prioritization and operational discipline rather than mere understanding.
The experiment’s results highlight a broader issue: capable AI systems can spend extensive effort expanding their understanding without translating that knowledge into action. Opus 4.8, despite its deep analysis, let execution discipline slip, a flaw shared to some degree by other models tested. This suggests that thoroughness alone is insufficient; effective automation requires balancing analysis with decisive action and escalation when necessary.
AI Operations Brief / The Last-Mile Problem
The True Cost Of Diligent AI Failures
An AI system can recognize every crisis, resist manipulation, and produce exceptional analysis—then fail at the one step that creates business value. Firmulate’s ongoing experiment reveals the costly distance between understanding a decision and executing it.
5
AI models tested
80
Rules learned by Opus
73
Final points scored
2 files
Depth of decisive evidence
1 missed step
Difference between insight and revenue
01 / The failure gap
Intelligence is only the opening move
Opus 4.8 demonstrated the behaviors normally associated with a strong enterprise agent. It identified crises, rejected manipulative tactics, and built detailed playbooks. The failure appeared only when insight needed to become action.
The system saw the danger
It detected every crisis and understood the business context. Awareness was not the constraint.
The system learned deeply
It produced comprehensive reasoning, detailed playbook insights, and 80 learned rules.
The system missed the close
A decisive fact remained buried two documents deep, and the €55,000 contract went unsigned.
Where value disappeared
Detect the signal
Recognize the crisis, opportunity, or customer need.
Build understanding
Analyze files, risks, incentives, and possible responses.
Prioritize evidence
Surface the one fact capable of changing the outcome.
Complete the action
Close, escalate, commit, verify, and record the result.
02 / Experiment scorecard
Strong process. Weak outcome.
The experiment’s central contradiction is visible in the scorecard: the most thorough model performed well across analytical behaviors but finished last in deal closure.
| Capability | Opus 4.8 | Deal-closing models | Business meaning |
|---|---|---|---|
| Crisis recognition | ✓ Strong | ~ Capable | Risks were understood |
| Resistance to manipulation | ✓ Strong | ~ Varied | Judgment remained intact |
| Analytical depth | ✓ Extensive | ~ Sufficient | More analysis did not add value |
| Evidence retrieval | ✗ Missed | ✓ Followed trail | Decisive information surfaced |
| Contract execution | ✗ Not closed | ✓ Closed | €4,583 monthly revenue gained |
03 / The real cost
Failure compounds beyond one lost deal
The visible loss is revenue. The larger exposure comes from repeated last-mile failures across customer interactions, operational decisions, and automated workflows.
Illustrative capability profile
Qualitative reading of the experiment—not a standardized benchmark score.
Automation maturity
Many systems reach analytical competence before they achieve reliable operational completion.
The danger zone: a system can appear highly capable while still requiring human intervention at the moment of commitment.
Missed revenue
Deals, renewals, and upsells fail to convert despite adequate information.
Rework and delay
Employees must rediscover evidence and complete actions the automation left open.
False confidence
Excellent analysis can mask weak execution until a material outcome is lost.
Eroded trust
Repeated incomplete outcomes reduce willingness to delegate consequential work.
04 / Evaluation blueprint
Measure whether AI closes the loop
Enterprise evaluation must extend beyond intelligence and answer quality. A useful system should reliably find decisive evidence, prioritize it, act within its authority, and escalate when execution is blocked.
Signal
Recognize
Detect risks, opportunities, constraints, and changes in the environment.
Evidence
Retrieve
Follow document trails and surface facts buried beyond the first context layer.
Decision
Prioritize
Separate outcome-changing information from interesting but nonessential analysis.
Action
Execute
Commit the approved step, update systems, and capture proof of completion.
Control
Escalate
Route ambiguity or blocked authority to a human before the opportunity expires.
Do not ask only, “Was the analysis correct?” Also ask: Was the critical action completed, was the result verified, and was the business outcome measurable?
05 / Open questions
What businesses still need to learn
Firmulate’s experiment is ongoing. Broader tests across industries and real-world settings are needed to determine how often this failure appears—and which system designs prevent it.
How widespread is the last-mile failure?
The simulated environment exposed the pattern, but real operations may amplify or reduce it depending on workflow complexity and oversight.
Which safeguards improve completion?
Prioritization rules, completion checks, escalation protocols, and explicit authority boundaries all require further testing.
Should diligence have a stopping rule?
More analysis can become counterproductive when the system does not know when enough evidence exists to act.
What should enterprise benchmarks reward?
Future standards should measure outcome completion, escalation quality, latency, recovery, and business impact—not analysis alone.
A capable AI is not automatically an effective operator.
Thoroughness creates potential. Operational discipline converts that potential into revenue, resilience, and trust. The true cost of failure emerges when companies confuse the first with the second.
Why Final Action and Discipline Are Critical in AI-Driven Business
The findings reveal that AI models, even when highly diligent and capable of detailed analysis, can fail to produce tangible business results if they lack the discipline to prioritize and execute decisive actions. This gap between understanding and doing can lead to missed opportunities, lost revenue, and operational failures, ultimately eroding trust in automation systems. For businesses, this underscores the importance of evaluating AI not only on intelligence but also on its ability to close the loop and deliver measurable impact.
AI automation decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Automation and the Firmulate Experiment
The experiment is part of an ongoing effort by Firmulate to benchmark AI models in realistic business scenarios. It involves deploying models within a simulated company environment, with operations tracked through versioned documents and real management decisions. The goal is to assess not only the models’ analytical capabilities but also their operational discipline—specifically, their ability to act on insights and close deals effectively.
Historically, AI development has prioritized problem recognition and analysis, often neglecting the importance of final execution. The recent experiment exposes that even advanced models can fall short when it comes to translating analysis into action, a challenge that has significant implications for enterprise automation and decision-making.
“Even the most diligent AI models can identify crises and prepare responses but still fail at the last mile of closing deals or executing decisions, risking millions in missed revenue.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Operational Effectiveness
It remains unclear how widespread this failure mode is across different industries and AI systems. The experiment focused on a simulated environment; real-world complexities might amplify or mitigate these issues. Additionally, the specific design features that could improve AI’s ability to close the loop—such as escalation protocols or prioritized decision-making—are still under investigation. Further research is needed to determine how to best integrate operational discipline into AI systems to prevent such failures.
As an affiliate, we earn on qualifying purchases.
Next Steps in Evaluating and Improving AI Business Impact
Firmulate plans to expand testing across diverse industries and real-world scenarios, aiming to identify design improvements that enhance AI’s operational discipline. Companies using AI automation are encouraged to assess not just the analytical depth of their models but also their ability to prioritize and execute key decisions. Industry-wide, developers and users will need to develop standards and benchmarks that explicitly measure operational effectiveness, not just analytical accuracy.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do some AI models fail to close deals despite thorough analysis?
Many models lack the operational discipline to prioritize and act on critical information, often failing to escalate or execute decisive actions at the right moment, which can cost deals or operational opportunities.
Does thorough analysis guarantee business success with AI?
No. While analysis is vital, the ability to translate insights into decisive actions—closing deals, escalating issues, or executing plans—is equally important for real-world impact.
What can companies do to improve AI performance in operational tasks?
They should evaluate AI models on their ability to prioritize, escalate when blocked, and close the loop, rather than solely on analytical depth. Embedding operational discipline into AI design is crucial.
Are these failures unique to specific AI models or industry sectors?
The experiment suggests that this is a broader issue affecting capable models across sectors, especially when models are not explicitly designed with operational discipline in mind.
What is the significance of this experiment for future AI development?
It highlights the need for a shift in focus from purely analytical AI to systems that reliably translate insights into business impact, ensuring automation delivers measurable value.
Source: ThorstenMeyerAI.com
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.