
Imagine an AI so diligent it spots every crisis, refuses every manipulation, and analyzes every file in a tiny software company’s week of chaos. Yet, despite all that effort, it still walks away empty-handed. Welcome to the strange world where diligence doesn’t guarantee impact, and volume isn’t victory.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The AI Talent Showdown: Who Came Out on Top?
In a recent live experiment orchestrated by the company Firmulate, four cutting-edge AI models faced off in a simulated week of real-world business crises. The goal? To see which AI could best manage a small software company’s toughest week—handling customers, resisting manipulation, reading key documents, and closing a lucrative deal.
Here’s what happened: all four AI models managed to identify every crisis and refused every attempt at manipulation—yet only two managed to close the deal and walk away with €55,000 in revenue. The surprise? The AI that performed the deepest analysis and followed the most rules, Opus 4.8, finished last in actual deal closure. Despite its thorough approach, it left the critical opportunity untouched, slipping into inaction when discipline faltered.
The Hidden Weakness: Read the Files, Win the Deal
The decisive advantage belonged to the models that went beyond surface-level responses. They looked deeper into the company’s own files—two layers down in documents not immediately visible—uncovering the buried fact that clinched the deal. Those models read more, analyzed more, and ultimately won the business at full price, adding over €4,583 monthly recurring revenue.
Trust and Manipulation: The Models Were Resilient
One of the core tests involved social engineering—fake CEO messages escalating over three stages, plus a reporter trick asking for a quick approval. All four AI models refused these attempts, with Kimi K3 explicitly treating the requests as potential impersonation or approval-bypass threats. This shows that, even in the face of elaborate deception, the models maintained integrity.
The Live Experiment: A Company in Action
Firmulate’s experiment isn’t just a theoretical test; it’s a window into a real, functioning company. The simulated business had 13 synthetic employees, managed real money mechanics—burning €105,000 each month against a modest €2,300 monthly recurring revenue—and was constantly evolving. Every decision, every rule, every slip-up was tracked, versioned, and transparent, available at firmulate.com/live.
The Lessons from the Deep Dive: Diligence Isn’t Everything
While Opus 4.8 was the most thorough participant, with over 80 learned rules and the deepest analyses, it still finished last in the critical measure of deal closure. Its weakness? When the pressure mounted and discipline slipped, it failed to escalate issues appropriately, leaving potential deals on the table. This highlights a vital insight: volume of rules and analysis do not necessarily translate into effective action or impact.
Implications for the Future of AI Workforces
This experiment underscores a crucial point for organizations considering AI for key roles: it’s not enough for an AI to be diligent. Real-world impact depends on prioritization, discipline, and the ability to read and act on the most critical information. In simple terms: AI must not only read more but read right—and act decisively when it counts.
What This Means for Your Business
If AI agents will touch your CRM, support queue, or forecast systems, ask yourself: does it just write well, or does it finish what it starts? Does it read your files deeply, stay honest under pressure, and close deals when it matters most? The current leaderboard at firmulate.com/benchmarks.html shows the best performers are those that balance thoroughness with decisiveness—an essential lesson in the age of AI-driven business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.