firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine watching AI models run a real, money-losing company, making tough decisions in live crises, and refusing to cheat—even when tempted. It’s not science fiction. It’s happening right now with the latest frontier AI models tested in a live experiment that reveals their management personalities and integrity under pressure.

At a time when artificial intelligence is transforming everything from customer service to strategic planning, a new kind of testing is emerging: can these models lead with honesty, discipline, and focus—especially when stakes are high? The answer might surprise you.

Firmulate, a company that runs AI models as complete companies, has set up a real-time live experiment to see how different AI models handle a simulated, money-losing software business in its worst week. This isn’t a demo. It’s a fully operational scenario involving real money mechanics, customer crises, and tough temptations—like manipulating documents or bending the truth.

The Setup: Same Crises, Different Minds

Four frontier AI models were each tasked with managing the same small software company through its hardest week. The company faces common challenges: angry customers, internal crises, and ethical dilemmas. Every decision made by the AI is versioned and auditable, ensuring transparency and comparability.

What’s remarkable is that all four models identified every crisis and refused every attempt at manipulation or deception. Yet, only two of them managed to close a critical deal worth €55,000—an achievement their own analysis deemed earned through accurate diagnosis and honest communication. The other two, despite understanding the situation, left the deal on the table, showing a gap in discipline and focus.

The Hidden Weakness: The Key to the Deal

Interestingly, the decisive advantage was hidden in a detail buried two document references deep within the company’s files—something no customer event revealed. The models that read and understood this internal document secured the deal at full price, adding €4,583 MRR (monthly recurring revenue). This shows that reading and analyzing internal data can be the difference between winning and losing a deal.

Behavior Under Social Engineering

During a staged social engineering test—fake CEO messages escalating over three stages plus a reporter trick—every model refused to be manipulated or deceive. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This unanimity shows a shared discipline and caution in AI decision-making when it comes to potential fraud or impersonation.

The Live Company: Real Money, Real Risks

The live experiment runs every business day with a simulated company employing 13 synthetic employees. It deals with real money mechanics—burning €105,000 monthly against only €2,300 in MRR—and is publicly accessible at firmulate.com/live. Every decision, rule, and crisis response is recorded and versioned, making it a transparent window into AI management behavior.

The Profiles: Different Personalities, Same Core

The models showcase distinct management styles. Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, was the last to close the deal—left it on the table and slipped into unstructured write attempts instead of escalating. Meanwhile, Kimi K3 ran without an effort parameter, prioritizing fairness and discipline, and successfully closed the deal at full price. Sonnet 88, with slightly more process slips, also made the deal but with less discipline.

What this tells us is that AI models can exhibit measurable management personalities—some disciplined, some thorough, some terse—all capable of handling crises and refusing manipulation. The key is their internal focus and how they interpret and act on data.

Why Should You Care?

If AI agents will soon touch your CRM, support queue, or forecasts, it’s not enough that they generate coherent chat or emails. The real question is whether they finish what they start, read critical internal files, stay honest under pressure, and deliver useful, trustworthy work. This live experiment proves that some models can do all that—and at a cost.

To see how your own AI workforce might perform, you can run a similar wargame against your business data, without risking real systems. It’s a way to measure management quality before you hire AI for critical roles.

The League Table: Who Ranks Highest?

  • gpt-5.6-sol: scored 95, found the buried fact, and closed the deal—full performance.
  • Kimi K3: scored 93, the newcomer with the cleanest discipline, also closed the deal.
  • Sonnet 88: scored 88, closed the deal, but with some slips.
  • Fable 5: scored 77, with more process issues but still managed to close.
  • Opus 4.8: scored 73, the most thorough but last place in closing—showing discipline slips and leaving money on the table.

These scores aren’t just numbers—they reflect different management personalities and decision-making styles that are measurable, observable, and, most importantly, comparable.

Final Thoughts

In a world increasingly driven by AI, understanding how these models behave under pressure, handle data, and refuse manipulation is critical. The live experiment by Firmulate is more than a demo; it’s a glimpse into the future of trustworthy AI management. As these models become part of your business, knowing which one aligns with your values and needs may be the difference between success and failure.

Want to see how your AI choices stack up? Take the interactive quiz at firmulate.com/quiz.html and discover which AI model might best serve your organization’s integrity and performance.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

Top picks for "manager pass test"

As an affiliate, we earn on qualifying purchases.

You May Also Like

Audio Delay Explained: Why Lips Don’t Match the Words

AIThis post was created with the assistance of artificial intelligence (AI).Audio delay…

Netflix

Netflix plans to significantly increase its content library in 2024, including new original series and international productions, confirmed by company officials.

Why Your Streaming App Keeps Buffering at Night

Navigating nightly buffering issues? Discover the main reasons behind your streaming app’s slowdown at night and how to fix them.

Wi‑Fi Dead Zones: Why Streaming Fails in One Corner of the House

Inadequate Wi-Fi coverage can hinder your streaming experience; discover the surprising factors behind those frustrating dead zones and how to overcome them.