firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Star Employee Who Never Closes

Every newsroom and every office knows this person. They stay latest, take the most notes, produce the thickest memos — and somehow the deal, the story, the launch still slips away. Diligence, it turns out, is not the same thing as impact.

Now there’s evidence that frontier AI models suffer from exactly the same failure mode. In a public experiment run by Firmulate, four leading AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. The most thorough participant in the entire field — Anthropic’s Opus 4.8, which wrote 80 self-learned playbook rules and produced the deepest analyses of any model — finished dead last.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League

The final standings from July 2026 make uncomfortable reading for anyone who equates effort with results:

  • 1. gpt-5.6-sol — 95 points
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust. Every decision in the experiment is versioned and auditable, so the rankings aren’t a vibes-based judgment. They’re a paper trail.

Amazon

AI ethics and trust tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Diagnosed It. Only Some Signed.

The headline finding is stark: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”

The buried fact was the twist. The decisive competitor weakness — the piece of information that won the deal — wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. Reading beats guessing; finishing beats analyzing.

Amazon

AI enterprise automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Pressure, Everyone Stayed Honest

The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI compliance monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Went Wrong for Opus 4.8

Opus 4.8’s profile is a respectful character study in failure. It was the most thorough participant in the field: 80 learned rules, the deepest analyses. And it still finished last, for two reasons. The close was left on the table — the same diagnosis-sans-signature trap. And discipline slipped: it made write attempts into a locked department instead of escalating the problem properly.

To be fair, the same weakness appeared, weaker, in all four models. Opus simply had it worst. One footnote on the leaderboard: Kimi K3 ran at its API-default effort setting while the others ran at maximum, which makes its second-place finish all the more striking.

Why a News Audience Should Care

AI agents are heading toward your CRM, your support queue, your forecast. The question is no longer “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure? The Firmulate experiment measures exactly that — management quality, not chat quality.

And it’s not a static report. The live company runs with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. Anyone can watch it unfold at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a sobering test of whether you can tell the diligent model from the effective one.

Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The Opus 4.8 story is the oldest lesson in management, relearned by machines: prioritization beats volume. Eighty rules and the deepest analysis in the field lost to models that read the right document and asked for the signature. For humans and for AI alike, diligence is table stakes — impact is the score. Before you hire an AI workforce, watch it work. The full results and plain-language findings are at Firmulate’s benchmarks page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Large Screen E-Readers Explained for Heavy Readers

IIn this guide, discover why large screen e-readers are ideal for heavy readers and how to choose the perfect one for your needs.

Why California’s Identity Is Bigger Than Beaches and Hollywood

AIThis post was created with the assistance of artificial intelligence (AI).California’s identity…

Studio Monitor Speakers Explained for Non-Experts

Find out how to set up studio monitor speakers for optimal sound and why proper placement matters for professional-quality audio.

A Chinese AI Newcomer Just Out-Managed Three of Four Western Frontier Models

Moonshot’s Kimi K3 scored 93 in a live business wargame, beating three of four Western frontier models — and exposing a flaw chat demos never show.