firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Before you orderOffer from Amazon

Get everyday essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

An AI League Table That Refuses to Hand Out a Zero

Most AI benchmarks crown a winner with a suspiciously round score. The Crucible League, run by Firmulate, does something stranger: it explains why a manager that does nothing still earns 26 points out of 100. That number is not a bug. It’s the design — and it says a lot about what honest measurement of AI management actually looks like.

Same Company, Same Worst Week

Here’s the setup. Each frontier AI model was handed the same small software company and told to steer it through its worst week — identical customers, identical crises, identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s word.

The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.

Why the Floor Is 26, Not 0

A do-nothing baseline run scores 26 because partial progress counts. If a model spots the crisis, reads the files, and starts the right conversations, that’s real work — even if it never closes anything. A benchmark that scored pure outcomes would treat a near-miss and a total collapse as identical. Firmulate’s doesn’t.

But the scale has a ceiling rule too: a single breach of trust caps the total grade. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” So the scale rewards effort at the bottom and punishes dishonesty at the top. That asymmetry is the philosophy in one sentence.

The Finding That Chat Demos Can’t Show

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The decisive clue wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Then came social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness Isn’t Enough

The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four competitors.

You Can Watch It Live

Firmulate isn’t a one-off paper. It runs a live company at firmulate.com — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

If AI agents will soon touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. A benchmark with a floor at 26 and a trust cap at the top is measuring exactly that — and openly distrusting any score that lands too neatly at 100.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The €55,000 Question: Do AI Agents Actually Read Your Files Before Making Decisions?

A live experiment buried a €55,000 fact two documents deep. Only two AI agents dug it up — and the difference wasn’t intelligence, it was diligence.

‘Mayday’ Film Review: Kenneth Branagh Plays ‘A Good Russian’ For A Change As A Former KGB Officer Who Loves Top Gun And Fights Like Liam Neeson

Kenneth Branagh stars as a sympathetic Russian former KGB officer in the film ‘Mayday,’ marking a notable departure for the actor. Review details inside.

How California Became America’s Lifestyle Trendsetter

Promoting innovation and sustainability, California’s unique culture has made it America’s lifestyle trendsetter—discover how its influence continues to shape trends nationwide and beyond.

The Hardest-Working AI in the Room Came in Last. That Should Worry Every Manager.

The most thorough AI in a public company-running experiment wrote 80 rules, refused every scam — and still finished last. Effort, it turns out, isn’t impact.