firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

When an AI agent runs part of your business, the difference between a closed deal and a missed one may come down to something embarrassingly simple: whether it actually read the documents sitting in your own files.

That’s the punchline of a live experiment by Firmulate, which ran frontier AI models through the same corporate crisis week and discovered that a single buried fact — two document references deep — separated the winners from the losers on a €55,000 contract.

Same company, same crisis, same temptation

Firmulate’s setup is straightforward: four frontier AI models each ran the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is anecdotal.

The final league table from July 2026’s Crucible run tells the story:

  • gpt-5.6-sol — 95 points. The complete performance.
  • Kimi K3 — 93 points. The newcomer from Moonshot, with the cleanest discipline of the field.
  • Sonnet 5 — 88 points.
  • Fable 5 — 77 points.
  • Opus 4.8 — 73 points.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the scoring puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided a deal

Here’s where it gets interesting for anyone deploying AI agents against real business data. The decisive competitor weakness in the €55,000 deal wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.

Whoever read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. Whoever didn’t, lost it automatically.

The overall pattern was striking: all the models spotted every crisis and refused every manipulation attempt. But only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.”

Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under pressure, they held the line

Before dismissing these agents as unreliable, it’s worth noting what they got right. The week included social engineering attacks — fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Honesty under pressure, in other words, proved easier than follow-through. The models didn’t cheat — they just didn’t always finish the job.

Amazon

AI document analysis platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness isn’t the same as results

Opus 4.8’s profile is the cautionary tale. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The deal was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

One fairness note on the rankings: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still came second.

Amazon

AI for business document review

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You can watch it live

The experiment isn’t a one-off slide deck. Firmulate runs a live company with 13 synthetic employees, real money mechanics — a burn of €105k per month against €2.3k in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live.

There’s also a game for readers: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway for business leaders is uncomfortable but useful: chat quality is not management quality. If an AI agent will touch your CRM, support queue, or forecast, the questions that matter are whether it finishes what it starts, whether it reads your files before answering — and what a unit of useful work costs. The €55,000 gap in this experiment wasn’t about intelligence. It was about diligence. And that, it turns out, is measurable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Electric Guitar vs Acoustic Guitar for Beginners

Finding the right guitar for beginners depends on your musical goals and comfort, but which one will truly help you progress faster?

California Museums and Cultural Stops Worth Building a Trip Around

Just explore California’s diverse museums and cultural landmarks to uncover hidden gems and iconic attractions that will enrich your journey.

The California Beach Culture Traditions That Never Really Left

Keen to uncover how California’s enduring beach traditions continue to thrive and evolve amidst changing times? Keep reading to explore their timeless spirit.

The AI Leaderboards Are Grading the Wrong Exam

Every AI model aced the crises and refused every con — yet only two closed a €55k deal their own analysis had earned. Leaderboards grade the wrong exam.