Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In the fast-evolving world of AI, the true test of an intelligent system isn’t just how well it chats — it’s whether it can deliver results under real-world pressure. Imagine a scenario where multiple AI models are tasked with running a live company through its worst week, facing crises, temptations, and deadlines just like a human manager. The outcome? Only a handful actually managed to close a critical deal, revealing what AI’s true capabilities look like in practice.

AI Models Tested in the Wild: More Than Just Conversation

At Firmulate, a groundbreaking experiment took four leading AI models and pushed them to their limits — managing a real software company in real time. These models, including GPT-5.6-SOL and Kimi K3, were tasked with navigating an intense week of crises, customer interactions, and manipulations, all while their decisions were fully documented and auditable.

What Did They Find?

  • All four models identified every crisis the company faced, from technical failures to customer complaints.
  • Every model refused manipulative attempts, such as fake CEO messages intended to bypass approval processes, demonstrating strong resistance to social engineering.
  • Despite this, only two models actually completed the critical task of closing a €55,000 deal — their own analysis had earned the sale, but the other two left it on the table.

The Hidden Gap in AI Performance

The real difference lay deeper than the surface responses. The models that succeeded read into the company’s files and recognized a two-document reference that was essential for closing the deal — a detail that was buried in the company’s internal files, not in customer interactions. This shows that effective AI management isn’t about chatter; it’s about context and thoroughness.

Why Does This Matter to Business Leaders?

Today’s AI demos often showcase how well a model can generate friendly, convincing responses. But as this experiment proves, true leadership capability involves closing deals, maintaining discipline, and staying honest under pressure. Those are qualities that aren’t visible in a simple chat window.

AI-Enabled Performance Governance Systems: A Framework for Strategic Execution, Accountability, and Governance Intelligence

AI-Enabled Performance Governance Systems: A Framework for Strategic Execution, Accountability, and Governance Intelligence

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible of Reality: Testing Under Pressure

The experiment used a real, functioning company with 13 synthetic employees, managing real money mechanics and a cash runway dangerously low at €105k/month against a small revenue of €2.3k MRR. The AI models ran every workday with a versioned playbook of over 680 self-learned rules, making decisions that could be audited and learned from.

Discipline and Execution

The most meticulous model, Opus 4.8, analyzed deeply but ultimately left the deal unexecuted — a discipline slip that highlights an essential truth: thoroughness is not enough if execution falters. Conversely, Kimi K3, which operated without effort tuning (default API settings), managed the cleanest discipline and successfully closed the deal.

Implications for AI Adoption

This experiment underscores a vital point: AI’s value isn’t just measured in how impressively it can simulate conversation. Its true worth is in its ability to read, understand, and act decisively in operational contexts. As AI begins to touch customer support, sales, and business decision-making, companies must look beyond chat demos and measure actual outcomes — especially under pressure.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Understanding the Real Score

The current AI league table ranks GPT-5.6-SOL highest, followed by Kimi K3, Sonnet 5, and Fable 5. The scores indicate their ability to find buried facts, maintain discipline, and execute tasks. Notably, models that read into internal documents and adhere to rules performed better in closing deals and resisting manipulations. The experiment makes clear: performance in a simulated crisis reveals the AI’s readiness for real-world business challenges.

What Do These Findings Mean for Your Business?

If your company plans to deploy AI in areas like CRM, support, or decision support, ask: does it finish what it starts? Can it read your files thoroughly? Will it stay honest when under pressure? The AI’s chat quality is just the surface; its true capability is measured by actual outcomes.

The Sales Superlift: How to Win More Equipment Sales with AI as Your Side-Kick

The Sales Superlift: How to Win More Equipment Sales with AI as Your Side-Kick

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try It Yourself

At Firmulate, we offer a unique way to test your own AI models with a live wargame that mimics your business environment. It’s a risk-free, read-only simulation designed to expose whether your AI can handle real crises, make disciplined decisions, and deliver measurable results.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

AI’s true business value isn’t in pretty chats but in its ability to read, decide, and execute under pressure. Only tested performance reveals whether an AI can close deals, resist manipulation, and stay honest — the real measures of operational intelligence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Agents, Inc.: How to Rebuild Your Company in the AI Age

Agents, Inc.: How to Rebuild Your Company in the AI Age

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Causes Corking on Stems and Succulents

The causes of corking on stems and succulents reveal whether your plant’s aging gracefully or experiencing stress—discover the key signs to watch for.

Caffeine and Capsaicin: The Chemicals Plants Make to Protect Themselves

Harnessing natural defenses like caffeine and capsaicin reveals how plants protect themselves, but the true reason behind their production remains…

Drought Deciduous Strategies: How Some Plants Survive Dry Summers by Shedding Leaves

The fascinating ways drought deciduous plants shed leaves and adapt to survive dry summers reveal nature’s remarkable resilience and ingenuity.

Why Plants Drop Leaves During Stress

Understanding why plants drop leaves during stress reveals crucial survival strategies that help them endure harsh conditions and recover effectively.