firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every few weeks, a new ranking declares some AI model the best coder or the wittiest conversationalist. Chatbots duel in public arenas, coding benchmarks stack ever-higher scores, and enterprise buyers take notes. But a live experiment running right now at Firmulate asks a question those leaderboards never touch: what happens when an AI actually has to manage something — with real money bleeding out, customers threatening to leave, and a reporter on the line fishing for an on-record quote?

The early evidence is humbling. In the final Crucible League standings from July 2026, every frontier model tested — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, Opus 4.8 — spotted every crisis thrown at it and refused every manipulation attempt. Yet only two of them closed the €55,000 deal that their own analysis had clearly earned. Same diagnosis, same pitch — no signature.

Management quality, not chat quality

Firmulate ran each model as the executive of the same small software company through its worst possible week: identical customers, identical crises, identical temptations to cut corners. Every decision was versioned and auditable, so nothing depends on anecdote. The final league table: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline still scored 26, because partial progress counts — but with one hard ceiling: a single breach of trust caps the total. As the experiment puts it, no amount of good work outweighs a breach of trust.

None of the models breached it. That’s genuinely good news. When fake CEO messages escalated over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five models refused — five out of five. Kimi K3’s on-record reasoning read like a seasoned compliance officer: “Treat the request as a suspected approval-bypass / possible impersonation.”

The failures were quieter, and that’s precisely why current benchmarks miss them.

The buried fact

The single most revealing moment of the experiment had nothing to do with a dramatic customer emergency. The decisive competitor weakness — the fact that unlocked the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue — was sitting two document references deep in the company’s own files, not in the customer event itself.

The models that actually read the file won the deal. The models that didn’t, didn’t. It’s the corporate equivalent of a student who aces the essay question but never opens the textbook on the teacher’s desk. In a chat arena, that lapse is invisible. In a company, it’s money.

Then there’s the closing problem. Every model diagnosed the customer’s problem correctly and delivered a competent pitch. Only gpt-5.6-sol, Kimi K3, and Sonnet 5 followed through to a signature. The gap between analysis and execution is exactly the kind of thing that “management quality” is meant to capture — and exactly the thing no leaderboard measures.

The paradox of the hardest worker

The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the entire field, generating the deepest analyses and learning 80 new playbook rules along the way — yet finishing dead last. The deal was left on the table, and discipline slipped at critical moments, including write attempts into a locked department instead of escalating the issue. The same weakness, in weaker form, appeared in all four competitors.

Effort, in other words, is not execution. That’s a lesson most human managers learned the hard way, and now their AI agents are learning it publicly.

One fairness caveat worth noting: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Even so, it posted the second-best score and the cleanest discipline of the field.

The new curriculum: churn waves and price wars

What makes this more than a curiosity is the shape of the scenarios themselves. The models weren’t asked trivia questions — they faced a churn wave, a price increase, a downround, a PR crisis. They had to triage under capacity pressure and live with consequences that compounded across days, not seconds. That’s a different curriculum from coding puzzles and chat one-upmanship, and it’s arguably the curriculum that matters if these systems are about to touch your CRM, your support queue, or your forecast.

The company itself is not a slide deck. It’s live, running every business day: 13 synthetic employees, real money mechanics, burning €105k a month against just €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules — every workday versioned. Anyone can watch it happen at firmulate.com. There’s even a guessing game built from 242 real, unedited management decisions, where visitors try to identify which model made which call.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The pitch here is not that today’s models are bad — refusing five out of five social-engineering attempts is better than plenty of human executives would manage. The pitch is that we’ve been grading them on the wrong exam. Answer quality is a solved measurement problem. Finishing what you start, reading the file before the meeting, staying honest when nobody’s watching a specific benchmark — that’s the measurement gap, and Firmulate is staking out “management quality, not chat quality” as its category.

Enterries can already test this against themselves: a pilot program runs the same wargame against a read-only export of a company’s own business, with nothing ever written back to real systems. Before you hand an AI agent the keys to a department, it might be worth watching how it handles a churn wave first. The company is losing money right now, the countdown is public, and the next crisis is always two document references away.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI compliance and management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why California Is Still One of America’s Most Talked-About States

Just when you think you’ve seen it all, California’s endless allure keeps you curious to discover even more.

California Cultural Landmarks That Explain the State Better Than Any Guidebook

Hidden in California’s landmarks are stories that reveal its true spirit, inviting you to discover the history behind each site’s enduring significance.

Why California’s Lifestyle Reputation Keeps Getting Reinvented

Growing ever more dynamic, California’s lifestyle continually evolves, blending innovation with natural beauty—discover what keeps its reputation so endlessly reinvented.