
Imagine a real company that has no human employees, runs 24/7 with synthetic workers, and is openly battling for survival. This is not science fiction — it’s the live experiment at firmulate.com/live.html. As you watch, you see the stark reality of artificial intelligence in business: decision-making under pressure, transparency in failures, and a relentless fight to stay afloat.
The Live Company: An Unusual Business Testbed
At the heart of this experiment is a small, virtual software company operated entirely by AI models. Every day, the company faces real crises — customer issues, ethical dilemmas, and strategic choices — just like a human-run business. Yet, what makes this setup extraordinary is that it is run by four different frontier AI models, each with its own decision-making style and weaknesses. The models are given the same challenging week of operations, with identical crises and temptations, and their decisions are meticulously recorded and analyzed.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decoding the AI Performance
In the latest benchmark, four AI models competed to manage this virtual company effectively. All four detected every crisis and refused every attempt at manipulation, such as social engineering tactics. However, their ability to secure the company’s most critical deal varied significantly. Two models successfully closed a €55,000 contract, but only one of them truly earned it through accurate diagnosis and proper pitch. The other, despite arriving at the same conclusion, left the deal on the table, revealing issues in discipline and decision process.

AI for Solo Lawyers: A Practical Guide to AI Tools that Save You Time and Grow Your Practice (AI for Professionals)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses in the Company’s Files
What’s particularly revealing is that the decisive advantage was buried two document references deep within the company’s files — information that only the models that read thoroughly could uncover. Those that missed this hidden detail failed to win the deal, illustrating that in AI decision-making, reading and understanding the full context can be the difference between success and failure. This underscores a core truth: attention to detail and comprehensive information processing are vital in high-stakes business decisions.

AUTONOMY WITHOUT CONTROL: What your AI agents can do that you can't undo
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Dealing with Ethical and Social Pressures
During simulated social engineering attacks, fake CEO messages and reporter tricks were used to test the models’ integrity. Remarkably, all five attempted manipulation scenarios failed uniformly. Kimi K3, one of the models, explicitly reasoned that such requests resembled impersonation or approval bypass attempts. This level of ethical resistance is crucial, especially as AI integrates more deeply into customer support, compliance, and decision frameworks.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of the Live Business
Beyond the decision-making tests, the company is a real, operational entity with 13 synthetic employees and actual money mechanics. It burns €105,000 each month while generating only €2,300 in monthly recurring revenue. It is openly running out of cash, with a public countdown to insolvency. Every workday, the company’s processes are versioned and publicly available for review, providing a brutal, transparent look at how AI behaves under real-world conditions.
Insights from the Performance Scores
The experiment’s scores reveal differing levels of performance:
- gpt-5.6-sol: Achieved the highest score of 95, successfully uncovering hidden facts and closing the deal — representing full performance.
- Kimi K3: Scored 93, managed to close the deal with the cleanest discipline, even without adjusting the effort parameters.
- Sonnet 5: Achieved 88, also closing the deal but with a few more process slips.
- Fable 5: Landed at 77, with more slips and missed opportunities, leaving potential value on the table.
This scoring system underscores that even among high-performing models, subtle differences in discipline, thoroughness, and decision process impact outcomes significantly.
What This Means for Business and AI
This ongoing experiment forces a stark question: in deploying AI for management or operational roles, is the focus on how well it writes or communicates, or on whether it truly completes tasks with integrity and thoroughness? The experiment demonstrates that the true test lies in whether AI can read full context, resist manipulation, and make decisions that hold under pressure — all critical for real-world business applications.
Accessible and Transparent
Interested readers can explore the full results and even participate in quizzes at firmulate.com/quotes.html. The live company runs every business day, and you can watch its decisions unfold or even run your own simulations against your business data, safely isolated from your real systems at firmulate.com/pilot.html.

This live experiment offers a rare window into how AI models perform under pressure in a real company setting — revealing strengths, weaknesses, and the vital qualities needed for AI to truly add value in business. It’s a high-stakes, transparent test that challenges assumptions about AI’s readiness for operational management and ethical decision-making.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html