firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Leaderboard Just Got Shaken

For most of the past two years, the story of frontier AI has been a two-horse race between American labs. The latest results from a live, publicly watchable experiment suggest that story is out of date. Moonshot’s Kimi K3 — a relative newcomer — finished second in a rigorous test of business management judgment, beating three of four Western frontier models and landing just two points behind the leader.

The test, run by Firmulate, isn’t a chatbot beauty contest. It puts AI models in charge of the same small software company during its worst week — same customers, same crises, same temptations — and scores how well they actually manage. In the final July 2026 league table, gpt-5.6-sol took first with 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 came in last at 73. For context, doing nothing at all scores 26.

Amazon

AI business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Crucible Actually Tests

Each model ran the identical company through identical trouble: a major sales opportunity, a security problem hidden in the company’s own files, a customer threatening to churn, and a stream of social engineering attacks. Every decision is versioned and auditable, so nothing rests on a judge’s impression.

The results cut against the usual AI narrative. All five models spotted every crisis, and all five refused every manipulation attempt — including a fake-CEO message campaign that escalated over three stages and a reporter’s seemingly harmless “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models closed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos, and it’s exactly the kind of thing that matters if an AI agent touches your CRM, support queue, or forecast.

Amazon

enterprise AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Needle

The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that did their homework before the pitch won the deal at full price, worth +€4,583 in monthly recurring revenue. Kimi K3 found the needle, closed the deal, saved the churning customer, and resisted all three baits with just a single deviation — the cleanest discipline in the field.

Amazon

AI cybersecurity threat detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

The most surprising profile belongs to Opus 4.8: it was the most thorough participant, adding 80 learned rules and producing the deepest analyses — yet finished last. It left the close on the table, and its discipline slipped, with write attempts into a locked department instead of escalating the issue. A weaker version of that same weakness — thorough analysis that doesn’t finish the job — appeared in all four other models.

Amazon

AI customer relationship management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch It Running

The backdrop is a live company, not a slide deck: 13 synthetic employees, real money mechanics — €105,000 monthly burn against €2,300 in MRR — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live. Curious readers can also try the “guess the model” quiz, built on 242 real, unedited management decisions, at firmulate.com/quiz.html. Full benchmark results are at firmulate.com/benchmarks.html, and enterprises can run the same wargame against a read-only export of their own business via the pilot program at firmulate.com/pilot.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The uncomfortable conclusion for buyers: model choice is no longer a safe assumption. A newcomer from Moonshot, running at its API-default effort setting, out-managed three of four Western frontier models on judgment, honesty under pressure, and follow-through. If you’re picking a model for real business work without running your own test, that’s now a bet — not a decision. As Firmulate puts it, the question isn’t whether an AI writes well. It’s whether it finishes what it starts, reads your files first, and stays honest when it matters.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh reasoning effort — a caveat worth keeping in mind when comparing scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Earthquake Of Magnitude 5.4 Strikes Egypt, Is Felt In Israel

A magnitude 5.4 earthquake struck Egypt and was felt in parts of Israel, causing minor damage and no reported injuries. Details are still emerging.

Electronic Drum Sets Explained for Smaller Spaces

An in-depth look at how electronic drum sets are ideal for small spaces, offering realistic sounds, quiet practice, and space-saving design—discover more below.

The AI Leaderboards Are Grading the Wrong Exam

Every AI model aced the crises and refused every con — yet only two closed a €55k deal their own analysis had earned. Leaderboards grade the wrong exam.

War Atlas: An Interactive Cartography Of Every Named War In Human History

A new online platform offers an interactive map charting every named war in human history, providing a comprehensive visual record for researchers and the public.