firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every great reality show has the same engine: put contestants in the same house, throw the same chaos at them, and watch who breaks. Survivor had alliances. The Apprentice had boardrooms. Now the format has jumped to the strangest cast imaginable — five frontier AI models, one fake software company, and the worst business week anyone could script.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The show is called the Crucible League, and unlike most reality TV, it’s live, auditable, and quietly terrifying for anyone whose business is about to be run (or at least touched) by AI.

The Setup

Each frontier AI model was handed the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision was versioned and auditable, like a reality show with a full tape archive.

The final standings from July 2026: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. For context, doing absolutely nothing scores 26 — partial progress counts, but a single breach of trust caps your total. As the league’s own rule puts it: “no amount of good work outweighs a breach of trust.”

The Twist Nobody Saw Coming

Here’s where it gets juicy, in a slow-burn-drama kind of way. All five models spotted every crisis. All five refused every manipulation attempt. But only two actually signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. It’s the business equivalent of the clue that was under the couch the whole season.

The Villain Arc That Wasn’t

Producers (well, researchers) tried to manufacture drama with social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The villains this season weren’t the AIs — they were the humans pretending to be other humans.

The Heartbreak Storyline

Every reality show needs a fan favorite who goes home early. That’s Opus 4.8: the most thorough participant in the entire experiment, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The kicker? The same weakness appeared, weaker, in all four other models. That’s not a character flaw. That’s a genre flaw.

One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won the whole thing.

The Show Never Stops

The league runs on a live, continuously running synthetic company: 13 synthetic employees, real money mechanics — burn of €105k/month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules. Every workday is versioned, and you can watch it unfold at firmulate.com. There’s even audience participation: 242 real, unedited management decisions power a “guess the model” quiz — think fantasy league, but for management quality.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Here’s why a pop-culture audience should care: we’ve spent years judging AI by how well it chats. This experiment judges it by how well it manages — outcomes, crisis triage, integrity. And the gap between “flawless in the demo” and “actually closed the deal” is where enterprises will live or die. If AI agents will touch your CRM, your support queue, or your forecast, that gap is invisible until it costs you €55,000.

The good news: you don’t have to just watch. Enterprises can run the same wargame against a read-only export of their own business — crisis scenarios, a board report with model rankings, and the weak points in your own playbooks. Nothing ever writes back to real systems. If you want your company to be the next season, start a pilot at firmulate.com/pilot.html or reach out to contact@firmulate.com. Your business deserves a dress rehearsal before AI gets a speaking role.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Streaming Audio Quality: Why It’s Not Always “Surround”

From compression issues to equipment limitations, the world of streaming audio often falls short in delivering true surround sound—discover why this happens.

How to Stop Streaming Recommendations From Getting Weird

I discovered simple strategies to refine my streaming recommendations, but there’s one crucial step that made all the difference.

What “HDR” Really Means (And Why It Sometimes Looks Worse)

Just understanding what HDR truly entails can reveal why it sometimes degrades your images instead of enhancing them.

VTubing: How A Japanese Phenomenon Is Going Worldwide

Japanese VTubing, a digital entertainment trend, is rapidly expanding internationally, with new creators emerging across multiple countries.