
Can AI Manage a Business Under Pressure? Watch It Live
Imagine observing a real company facing a week of relentless crises—customer complaints, internal dilemmas, and the temptation to cheat. Now, picture this company run not by humans, but by AI models that are analyzed and tested publicly, every decision recorded and scrutinized. This is the premise behind Firmulate, a groundbreaking live experiment where AI models operate as tiny, simulated companies in real-time, revealing how management quality could be measured and improved in ways never before possible.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Public Company, the Tiny Workforce, and the High Stakes
Firmulate’s experiment is straightforward but radical: four leading AI models are tasked with running a small software company through its worst week. The company has 13 synthetic employees, real money mechanics burning through €105,000 each month against a modest €2,300 monthly recurring revenue. Every workday is versioned and made auditable, giving viewers a transparent window into how each AI handles crises, customer demands, and ethical temptations.
How Do the Models Perform?
All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—successfully identified every crisis and refused all manipulation attempts, including a staged social engineering attack where fake CEO messages and a reporter’s trick were used to try and bypass controls. The models demonstrated remarkable discipline, with Kimi K3 earning the highest score of 93, and gpt-5.6-sol following closely at 95. Significantly, only two models managed to close the deal, signing a €55,000 contract that their own analysis had earned, highlighting the difference between diagnosis and execution.
The Hidden Vulnerability: The Critical Document
Delving into the company’s internal files, the experiment uncovered a key weakness—hidden two document references deep—that the models which read the file at depth were able to exploit. These models secured a deal at full price, adding over €4,583 monthly recurring revenue, illustrating how deep document analysis can be a decisive advantage.
Lessons on Integrity and Decision-Making
Throughout the experiment, every decision was transparent and auditable. The models faced social engineering attempts and rightly refused to sign off on non-verified requests, aligning with Kimi K3’s stated reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined behavior underscores a crucial point—AI can be trained to uphold ethical standards even when under pressure, something critical for real-world management systems.

Interview with the MONSTER AI: A Conversation about Power, Truth, and the Future of Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Burn Rate and Public Scrutiny
This isn’t a hypothetical scenario: the live-sided company is real, with a public cash countdown and a daily routine that is openly accessible at firmulate.com/live. It runs every workday, versioned and open for viewers to see how decisions unfold under the weight of economic and ethical pressures. The model-driven management costs €105,000 per month to operate, versus its modest €2,300 income—highlighting the urgent need for more disciplined and trustworthy AI management tools.
What Does This Teach Developers and Managers?
For software developers, QA specialists, and business leaders, the experiment offers a stark lesson: the true test of AI’s usefulness isn’t how well it chats or generates code, but whether it can finish what it starts, read critical documents, and remain honest under pressure. The models performed well on crisis detection and refusal of manipulation—but the real-world application depends on consistent execution and discipline, as seen in the case of Opus 4.8, which, despite thorough analysis, left a critical deal unexecuted due to process slips.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Building in Public — Transparency as a Management Tool
Firmulate’s approach is transparency taken to the extreme. Every decision, every rule learned, every crisis response is visible and auditable. The platform also hosts a quiz where 242 real management decisions are used to challenge viewers to guess which AI model made what decision, fostering understanding of AI decision-making and discipline.
Practical Applications and Future Outlook
Enterprises can leverage this live wargame—called a pilot—to test their own business models against AI decision-making. Nothing ever writes back to real systems, so companies can simulate and improve their AI-driven management strategies risk-free. The ongoing experiments continue to grow, with new benchmark runs queued and published automatically, fostering a rich, evolving understanding of AI leadership capabilities.


Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Takeaways
Observing AI managing a tiny company live reveals that, while models can identify crises and refuse manipulation, consistent execution and ethical discipline remain challenging. Deep document analysis can be a crucial advantage, and transparency in decision-making helps build trust. For businesses considering AI management systems, the experiment underscores an essential truth: trustworthiness and discipline are as vital as intelligence. Watch the experiment unfold at firmulate.com/live and learn what the future of AI-managed companies might look like.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html