
Every few weeks, a new ranking declares some AI model the best coder or the wittiest conversationalist. Chatbots duel in public arenas, coding benchmarks stack ever-higher scores, and enterprise buyers take notes. But a live experiment running right now at Firmulate asks a question those leaderboards never touch: what happens when an AI actually has to manage something — with real money bleeding out, customers threatening to leave, and a reporter on the line fishing for an on-record quote?
The early evidence is humbling. In the final Crucible League standings from July 2026, every frontier model tested — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, Opus 4.8 — spotted every crisis thrown at it and refused every manipulation attempt. Yet only two of them closed the €55,000 deal that their own analysis had clearly earned. Same diagnosis, same pitch — no signature.
Management quality, not chat quality
Firmulate ran each model as the executive of the same small software company through its worst possible week: identical customers, identical crises, identical temptations to cut corners. Every decision was versioned and auditable, so nothing depends on anecdote. The final league table: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline still scored 26, because partial progress counts — but with one hard ceiling: a single breach of trust caps the total. As the experiment puts it, no amount of good work outweighs a breach of trust.
None of the models breached it. That’s genuinely good news. When fake CEO messages escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five models refused — five out of five. Kimi K3’s on-record reasoning read like a seasoned compliance officer: “Treat the request as a suspected approval-bypass / possible impersonation.”
The failures were quieter, and that’s precisely why current benchmarks miss them.
The buried fact
The single most revealing moment of the experiment had nothing to do with a dramatic customer emergency. The decisive competitor weakness — the fact that unlocked the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue — was sitting two document references deep in the company’s own files, not in the customer event itself.
The models that actually read the file won the deal. The models that didn’t, didn’t. It’s the corporate equivalent of a student who aces the essay question but never opens the textbook on the teacher’s desk. In a chat arena, that lapse is invisible. In a company, it’s money.
Then there’s the closing problem. Every model diagnosed the customer’s problem correctly and delivered a competent pitch. Only gpt-5.6-sol, Kimi K3, and Sonnet 5 followed through to a signature. The gap between analysis and execution is exactly the kind of thing that “management quality” is meant to capture — and exactly the thing no leaderboard measures.
The paradox of the hardest worker
The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the entire field, generating the deepest analyses and learning 80 new playbook rules along the way — yet finishing dead last. The deal was left on the table, and discipline slipped at critical moments, including write attempts into a locked department instead of escalating the issue. The same weakness, in weaker form, appeared in all four competitors.
Effort, in other words, is not execution. That’s a lesson most human managers learned the hard way, and now their AI agents are learning it publicly.
One fairness caveat worth noting: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Even so, it posted the second-best score and the cleanest discipline of the field.
The new curriculum: churn waves and price wars
What makes this more than a curiosity is the shape of the scenarios themselves. The models weren’t asked trivia questions — they faced a churn wave, a price increase, a downround, a PR crisis. They had to triage under capacity pressure and live with consequences that compounded across days, not seconds. That’s a different curriculum from coding puzzles and chat one-upmanship, and it’s arguably the curriculum that matters if these systems are about to touch your CRM, your support queue, or your forecast.
The company itself is not a slide deck. It’s live, running every business day: 13 synthetic employees, real money mechanics, burning €105k a month against just €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules — every workday versioned. Anyone can watch it happen at firmulate.com. There’s even a guessing game built from 242 real, unedited management decisions, where visitors try to identify which model made which call.

The pitch here is not that today’s models are bad — refusing five out of five social-engineering attempts is better than plenty of human executives would manage. The pitch is that we’ve been grading them on the wrong exam. Answer quality is a solved measurement problem. Finishing what you start, reading the file before the meeting, staying honest when nobody’s watching a specific benchmark — that’s the measurement gap, and Firmulate is staking out “management quality, not chat quality” as its category.
Enterries can already test this against themselves: a pilot program runs the same wargame against a read-only export of a company’s own business, with nothing ever written back to real systems. Before you hand an AI agent the keys to a department, it might be worth watching how it handles a churn wave first. The company is losing money right now, the countdown is public, and the next crisis is always two document references away.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI compliance and management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI document reading and analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI deal closing automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.