AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In a world increasingly driven by AI, the true measure of a machine’s value isn’t just speed or volume — it’s whether it can prioritize, read carefully, and stay honest when it counts. A recent live experiment with four top AI models exposed a stark reality: even the most thorough and diligent AI can falter if it neglects to focus on what matters most. The lesson? Diligence alone isn’t enough; focus and discipline are key.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Benchmark: An AI Showdown in a Simulated Business Crisis

At the heart of this experiment was the Crucible League, an annual contest where leading AI models are tested through a simulated week in a small software company. The models faced identical crises: customer issues, ethical dilemmas, and manipulative attempts designed to tempt them into shortcuts or breaches of trust. The aim was simple: could these AIs manage the chaos and close a critical deal worth over €55,000 in recurring revenue?

The four models tested were GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8. Each ran the same scenario, with decisions fully versioned and auditable. The results? All four identified every crisis and refused every manipulation attempt—an encouraging sign of integrity. But when it came time to close the deal, only two models succeeded: GPT-5.6-sol and Kimi K3.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading the Files That Held the Key

The decisive factor wasn’t just crisis management but the AI’s ability to uncover the company’s buried facts. The critical piece of information that sealed the deal was stored two document references deep in the company’s files, not in the immediate customer interactions. Only the models that thoroughly read and analyzed these files managed to close at full price, adding over €4,583 MRR in value. The others overlooked this crucial insight, leaving the opportunity on the table.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Challenge of Disciplined Focus

One standout participant was Opus 4.8, praised for its thoroughness — it learned over 80 rules and conducted deep analyses. Yet, it ultimately finished last in the final score because it let discipline slip: instead of escalating critical findings, it wrote attempts into a locked department, ignoring their importance. This lapse showed that even a model with comprehensive rules can falter if it doesn’t prioritize the right information.

Amazon

AI business decision support system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ethical Tests: Refusing Manipulation and Social Engineering

Beyond operational decision-making, the models faced social engineering attempts involving fake CEO messages and a reporter trick. All five models refused to engage, citing concerns like impersonation or bypassing approval processes. Kimi K3 clarified: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrated that even the most advanced models understand and uphold ethical boundaries when pressured.

Amazon

AI focus and prioritization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business: A Company in the Wild

The experiment isn’t just theoretical. It’s run on a live, functioning company with 13 synthetic employees, managing real cash flow of €2.3k MRR against burn rates of €105k per month. Every day, the models make decisions that impact the business, which is publicly observable at firmulate.com/live. This transparency underscores the importance of testing AI in realistic, high-stakes environments rather than just chat demos.

The Bigger Takeaway: Prioritization Over Volume

Despite the depth of analysis, the experiment reveals a crucial insight: diligence alone isn’t enough. The most thorough participant, Opus 4.8, still lost because it failed to escalate critical issues, illustrating that focus and discipline are paramount. The same weakness appeared, albeit less strongly, in all models. Prioritizing the most important information — especially hidden insights — makes the difference between closing a deal at full price or leaving it on the table.

What Business Leaders Should Take Away

As AI increasingly touches facets of management, support, and decision-making, the question isn’t just whether the AI writes well. It’s whether it stays honest, reads deeply enough, and prioritizes what truly matters under pressure. The experiment shows that even the most capable AI can underperform if it lacks discipline and focus, highlighting the need for careful testing before deploying these systems in critical roles.

Practical Tools: Wargaming Your AI Workforce

To mitigate risks, firms can run their own “wargame” scenarios against AI models, using tools like the live platform at firmulate.com/pilot.html. These simulations replicate real crises without ever changing actual systems, giving managers a clear view of how AI behaves in high-stakes situations. This approach allows enterprises to identify weaknesses and build disciplined AI teams that excel not just in analysis but in decisive, honest action.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Plastic Plants? How Microplastics in Soil Affect Plant Growth

Unearthing the impact of microplastics in soil reveals how these tiny pollutants threaten plant growth and ecosystem health, leaving critical effects yet to be fully understood.

Sunburnt: How UV Light Damages Plants and How They Cope

The truth about UV damage to plants reveals surprising survival strategies that may help your garden thrive despite sun exposure.

Is The Economist Always Wrong?

An analysis of The Economist’s forecasting record, examining claims of consistent errors and the implications for its credibility.

Botanical Revival: Bringing Back Extinct Plants Through Seed Banks and Breeding

Never underestimate how seed banks and breeding can revive extinct plants, but the full story behind botanical revival is even more fascinating.