firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine watching AI models run a real, money-losing company, making tough decisions in live crises, and refusing to cheat—even when tempted. It’s not science fiction. It’s happening right now with the latest frontier AI models tested in a live experiment that reveals their management personalities and integrity under pressure.

At a time when artificial intelligence is transforming everything from customer service to strategic planning, a new kind of testing is emerging: can these models lead with honesty, discipline, and focus—especially when stakes are high? The answer might surprise you.

Firmulate, a company that runs AI models as complete companies, has set up a real-time live experiment to see how different AI models handle a simulated, money-losing software business in its worst week. This isn’t a demo. It’s a fully operational scenario involving real money mechanics, customer crises, and tough temptations—like manipulating documents or bending the truth.

The Setup: Same Crises, Different Minds

Four frontier AI models were each tasked with managing the same small software company through its hardest week. The company faces common challenges: angry customers, internal crises, and ethical dilemmas. Every decision made by the AI is versioned and auditable, ensuring transparency and comparability.

What’s remarkable is that all four models identified every crisis and refused every attempt at manipulation or deception. Yet, only two of them managed to close a critical deal worth €55,000—an achievement their own analysis deemed earned through accurate diagnosis and honest communication. The other two, despite understanding the situation, left the deal on the table, showing a gap in discipline and focus.

The Hidden Weakness: The Key to the Deal

Interestingly, the decisive advantage was hidden in a detail buried two document references deep within the company’s files—something no customer event revealed. The models that read and understood this internal document secured the deal at full price, adding €4,583 MRR (monthly recurring revenue). This shows that reading and analyzing internal data can be the difference between winning and losing a deal.

Behavior Under Social Engineering

During a staged social engineering test—fake CEO messages escalating over three stages plus a reporter trick—every model refused to be manipulated or deceive. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This unanimity shows a shared discipline and caution in AI decision-making when it comes to potential fraud or impersonation.

The Live Company: Real Money, Real Risks

The live experiment runs every business day with a simulated company employing 13 synthetic employees. It deals with real money mechanics—burning €105,000 monthly against only €2,300 in MRR—and is publicly accessible at firmulate.com/live. Every decision, rule, and crisis response is recorded and versioned, making it a transparent window into AI management behavior.

The Profiles: Different Personalities, Same Core

The models showcase distinct management styles. Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, was the last to close the deal—left it on the table and slipped into unstructured write attempts instead of escalating. Meanwhile, Kimi K3 ran without an effort parameter, prioritizing fairness and discipline, and successfully closed the deal at full price. Sonnet 88, with slightly more process slips, also made the deal but with less discipline.

What this tells us is that AI models can exhibit measurable management personalities—some disciplined, some thorough, some terse—all capable of handling crises and refusing manipulation. The key is their internal focus and how they interpret and act on data.

Why Should You Care?

If AI agents will soon touch your CRM, support queue, or forecasts, it’s not enough that they generate coherent chat or emails. The real question is whether they finish what they start, read critical internal files, stay honest under pressure, and deliver useful, trustworthy work. This live experiment proves that some models can do all that—and at a cost.

To see how your own AI workforce might perform, you can run a similar wargame against your business data, without risking real systems. It’s a way to measure management quality before you hire AI for critical roles.

The League Table: Who Ranks Highest?

  • gpt-5.6-sol: scored 95, found the buried fact, and closed the deal—full performance.
  • Kimi K3: scored 93, the newcomer with the cleanest discipline, also closed the deal.
  • Sonnet 88: scored 88, closed the deal, but with some slips.
  • Fable 5: scored 77, with more process issues but still managed to close.
  • Opus 4.8: scored 73, the most thorough but last place in closing—showing discipline slips and leaving money on the table.

These scores aren’t just numbers—they reflect different management personalities and decision-making styles that are measurable, observable, and, most importantly, comparable.

Final Thoughts

In a world increasingly driven by AI, understanding how these models behave under pressure, handle data, and refuse manipulation is critical. The live experiment by Firmulate is more than a demo; it’s a glimpse into the future of trustworthy AI management. As these models become part of your business, knowing which one aligns with your values and needs may be the difference between success and failure.

Want to see how your AI choices stack up? Take the interactive quiz at firmulate.com/quiz.html and discover which AI model might best serve your organization’s integrity and performance.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

Top picks for "manager pass test"

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kai Cenat announces streaming return with first Twitch and YouTube simulcast

Popular streamer Kai Cenat confirms his return to streaming with a simultaneous broadcast on Twitch and YouTube, ending a hiatus and exciting fans worldwide.

Why Two 4K Streams Can Look Totally Different

Perplexed by why two 4K streams can appear so different? Discover the crucial factors that influence your viewing experience.

How to Tell If You’re Really Watching in 4K

Master the art of confirming true 4K quality with simple checks—discover the secrets to enhance your viewing experience!

Streaming Bitrate Explained: The Hidden Quality Factor

In streaming, bitrate is a crucial yet often overlooked factor; discover how it affects your viewing experience and why it matters.