
Get everyday essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
An AI League Table That Refuses to Hand Out a Zero
Most AI benchmarks crown a winner with a suspiciously round score. The Crucible League, run by Firmulate, does something stranger: it explains why a manager that does nothing still earns 26 points out of 100. That number is not a bug. It’s the design — and it says a lot about what honest measurement of AI management actually looks like.
Same Company, Same Worst Week
Here’s the setup. Each frontier AI model was handed the same small software company and told to steer it through its worst week — identical customers, identical crises, identical temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anyone’s word.
The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.
Why the Floor Is 26, Not 0
A do-nothing baseline run scores 26 because partial progress counts. If a model spots the crisis, reads the files, and starts the right conversations, that’s real work — even if it never closes anything. A benchmark that scored pure outcomes would treat a near-miss and a total collapse as identical. Firmulate’s doesn’t.
But the scale has a ceiling rule too: a single breach of trust caps the total grade. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” So the scale rewards effort at the bottom and punishes dishonesty at the top. That asymmetry is the philosophy in one sentence.
The Finding That Chat Demos Can’t Show
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The decisive clue wasn’t in the customer conversation at all: it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
Then came social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness Isn’t Enough
The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four competitors.
You Can Watch It Live
Firmulate isn’t a one-off paper. It runs a live company at firmulate.com — 13 synthetic employees, real money mechanics, a burn of €105k a month against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
If AI agents will soon touch your CRM, your support queue, or your forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files before acting, and stay honest under pressure. A benchmark with a floor at 26 and a trust cap at the top is measuring exactly that — and openly distrusting any score that lands too neatly at 100.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
