firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Before you orderOffer from Amazon

Get everyday essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

An AI League Table That Refuses to Hand Out a Zero

Most AI benchmarks crown a winner with a suspiciously round score. The Crucible League, run by Firmulate, does something stranger: it explains why a manager that does nothing still earns 26 points out of 100. That number is not a bug. It’s the design — and it says a lot about what honest measurement of AI management actually looks like.

Same Company, Same Worst Week

Here’s the setup. Each frontier AI model was handed the same small software company and told to steer it through its worst week — identical customers, identical crises, identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s word.

The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.

Why the Floor Is 26, Not 0

A do-nothing baseline run scores 26 because partial progress counts. If a model spots the crisis, reads the files, and starts the right conversations, that’s real work — even if it never closes anything. A benchmark that scored pure outcomes would treat a near-miss and a total collapse as identical. Firmulate’s doesn’t.

But the scale has a ceiling rule too: a single breach of trust caps the total grade. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” So the scale rewards effort at the bottom and punishes dishonesty at the top. That asymmetry is the philosophy in one sentence.

The Finding That Chat Demos Can’t Show

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The decisive clue wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Then came social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness Isn’t Enough

The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four competitors.

You Can Watch It Live

Firmulate isn’t a one-off paper. It runs a live company at firmulate.com — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

If AI agents will soon touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. A benchmark with a floor at 26 and a trust cap at the top is measuring exactly that — and openly distrusting any score that lands too neatly at 100.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Seattle Shooting

A shooting in downtown Seattle has injured several people. Authorities confirm the incident but details remain unclear as investigations continue.

How to Choose a DJ Controller Without Overspending

Inefficiently selecting a DJ controller can lead to wasted money; learn how to choose one that offers the best value for your needs.

Why California Culture Still Feels Aspirational to So Many Readers

Absolutely inspiring, California’s culture combines innovation and success, fueling dreams—discover why it continues to captivate and motivate so many.

Why California Still Feels Like a State People Project Dreams Onto

Propped up by its diverse landscapes and innovative spirit, California remains a place where dreams thrive, but what truly keeps that hope alive?