
When an AI agent runs part of your business, the difference between a closed deal and a missed one may come down to something embarrassingly simple: whether it actually read the documents sitting in your own files.
That’s the punchline of a live experiment by Firmulate, which ran frontier AI models through the same corporate crisis week and discovered that a single buried fact — two document references deep — separated the winners from the losers on a €55,000 contract.
Same company, same crisis, same temptation
Firmulate’s setup is straightforward: four frontier AI models each ran the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is anecdotal.
The final league table from July 2026’s Crucible run tells the story:
- gpt-5.6-sol — 95 points. The complete performance.
- Kimi K3 — 93 points. The newcomer from Moonshot, with the cleanest discipline of the field.
- Sonnet 5 — 88 points.
- Fable 5 — 77 points.
- Opus 4.8 — 73 points.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the scoring puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The buried fact that decided a deal
Here’s where it gets interesting for anyone deploying AI agents against real business data. The decisive competitor weakness in the €55,000 deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.
Whoever read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Whoever didn’t, lost it automatically.
The overall pattern was striking: all the models spotted every crisis and refused every manipulation attempt. But only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.”
enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Under pressure, they held the line
Before dismissing these agents as unreliable, it’s worth noting what they got right. The week included social engineering attacks — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Honesty under pressure, in other words, proved easier than follow-through. The models didn’t cheat — they just didn’t always finish the job.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as results
Opus 4.8’s profile is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The deal was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
One fairness note on the rankings: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still came second.
As an affiliate, we earn on qualifying purchases.
You can watch it live
The experiment isn’t a one-off slide deck. Firmulate runs a live company with 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live.
There’s also a game for readers: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The takeaway for business leaders is uncomfortable but useful: chat quality is not management quality. If an AI agent will touch your CRM, support queue, or forecast, the questions that matter are whether it finishes what it starts, whether it reads your files before answering — and what a unit of useful work costs. The €55,000 gap in this experiment wasn’t about intelligence. It was about diligence. And that, it turns out, is measurable.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html