
Get everyday essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Star Employee Who Never Closes
Every newsroom and every office knows this person. They stay latest, take the most notes, produce the thickest memos — and somehow the deal, the story, the launch still slips away. Diligence, it turns out, is not the same thing as impact.
Now there’s evidence that frontier AI models suffer from exactly the same failure mode. In a public experiment run by Firmulate, four leading AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. The most thorough participant in the entire field — Anthropic’s Opus 4.8, which wrote 80 self-learned playbook rules and produced the deepest analyses of any model — finished dead last.
As an affiliate, we earn on qualifying purchases.
The Crucible League
The final standings from July 2026 make uncomfortable reading for anyone who equates effort with results:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust. Every decision in the experiment is versioned and auditable, so the rankings aren’t a vibes-based judgment. They’re a paper trail.
As an affiliate, we earn on qualifying purchases.
Everyone Diagnosed It. Only Some Signed.
The headline finding is stark: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”
The buried fact was the twist. The decisive competitor weakness — the piece of information that won the deal — wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. Reading beats guessing; finishing beats analyzing.
As an affiliate, we earn on qualifying purchases.
Under Pressure, Everyone Stayed Honest
The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI compliance monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Went Wrong for Opus 4.8
Opus 4.8’s profile is a respectful character study in failure. It was the most thorough participant in the field: 80 learned rules, the deepest analyses. And it still finished last, for two reasons. The close was left on the table — the same diagnosis-sans-signature trap. And discipline slipped: it made write attempts into a locked department instead of escalating the problem properly.
To be fair, the same weakness appeared, weaker, in all four models. Opus simply had it worst. One footnote on the leaderboard: Kimi K3 ran at its API-default effort setting while the others ran at maximum, which makes its second-place finish all the more striking.
Why a News Audience Should Care
AI agents are heading toward your CRM, your support queue, your forecast. The question is no longer “does it write well.” It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure? The Firmulate experiment measures exactly that — management quality, not chat quality.
And it’s not a static report. The live company runs with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. Anyone can watch it unfold at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions — a sobering test of whether you can tell the diligent model from the effective one.
Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Takeaway
The Opus 4.8 story is the oldest lesson in management, relearned by machines: prioritization beats volume. Eighty rules and the deepest analysis in the field lost to models that read the right document and asked for the signature. For humans and for AI alike, diligence is table stakes — impact is the score. Before you hire an AI workforce, watch it work. The full results and plain-language findings are at Firmulate’s benchmarks page.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
