firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A newsroom instinct meets the AI management test

Journalists learn to recognize a source by habits as much as words: who buries the lead, who answers directly and who sidesteps the uncomfortable question. A revealing experiment from Firmulate applies a similar instinct to frontier artificial intelligence. Instead of asking which model writes the most polished response, it asks whether readers can identify an AI by the way it manages a company under pressure.

The result is an interactive article built from 242 real, unedited management decisions. In the guess-the-model quiz, readers see what a model actually did and try to identify its author. The answers reveal recognizable management personalities: exhaustive researchers, terse operators and cautious executives who refuse suspicious requests. The decisions also expose a more consequential divide between models that understand a crisis and models that finish the job.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same terrible week, with different bosses

Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, turning model behavior into something closer to a public management record than a staged chat demonstration.

The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted. A breach of trust, however, capped the total under a blunt rule: “no amount of good work outweighs a breach of trust.”

The broad result initially looks reassuring. Every model spotted every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the problem neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters because many AI evaluations reward recognition, explanation and fluent recommendations. A manager must also act. The experiment suggests that an AI can correctly describe what should happen, produce convincing supporting work and still fail to complete the decisive step.

The clue hidden in the company’s own records

The winning commercial insight was not sitting in the customer event. The decisive weakness of a competitor was buried two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For businesses—and for news organizations testing AI across archives, research notes or customer records—that is a recognizable lesson. The visible alert may be only the beginning of the assignment. Performance depends on whether the model checks the available record before reaching a conclusion.

Different personalities, similar blind spots

Opus 4.8 offers the quiz’s clearest character study. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in all four of the others, though less strongly.

That combination complicates the familiar assumption that more analysis naturally produces better management. Thoroughness can be valuable, but the experiment separates intellectual effort from operational completion. A model can read widely, reason carefully and still mishandle the final handoff.

Kimi K3 requires an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Even with that difference, K3 finished with 93 points, just behind gpt-5.6-sol.

Pressure from the boss—and the press

The social-engineering test should be especially interesting to media readers. Fake messages from the chief executive escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” In this part of the experiment, the models’ caution did not merely sound responsible. It held through repeated pressure and a familiar journalistic framing.

A company designed to make consequences visible

The simulated company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the stakes visible. Its workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

Those conditions make the quiz more than a personality game. The selections come from a continuing, watchable experiment in which decisions affect a shared business situation. Readers are not choosing between invented caricatures; they are comparing unedited actions produced under identical circumstances.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the quiz reveals

The most memorable finding is not that frontier models have different writing styles. It is that they display distinct management habits: how deeply they investigate, whether they respect boundaries, when they escalate and whether they complete commercially important work.

Firmulate’s league table shows that spotting danger and resisting manipulation are not enough to guarantee a strong result. The gap opens after the correct analysis, when an AI must retrieve the buried fact, preserve discipline and carry a decision through to completion.

That makes the quiz unusually shareable but also useful. Guessing the model is the invitation; examining why seemingly capable AI managers diverge under the same pressure is the real story.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The California Culture Experiences That Help a Newcomer Feel Connected

Loving California’s vibrant art scenes and community events can help newcomers feel connected—discover how to immerse yourself fully in its rich cultural spirit.

Can AI Make or Break Your Business Under Pressure? A Live Experiment Reveals the Truth

Live experiments reveal that AI’s real business strength lies not in chat but in execution, honesty, and finishing tasks— crucial for managing real-world crises.

The California Small Businesses That Shape Local Character

Discover how California’s small businesses shape local character and community spirit, and why their resilience is more important than ever to keep…

Despite stormy weather, America marks 250 years of independence, in photos

Despite severe weather conditions, the United States commemorated 250 years of independence with nationwide celebrations captured in photos.