AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Recipe You Didn’t Taste Is the Recipe You Can’t Trust

Any cook who has hosted a dinner party knows the rule: you never serve a dish you haven’t tasted. A sauce can look glossy and smell perfect and still be under-seasoned where it matters. The same is true — it turns out — for AI models. They can demo beautifully in chat, answer fluently, impress in a tasting spoon. But can they run a kitchen through a Friday-night rush without dropping a plate, burning the pan, or quietly pocketing the tips?

That question is exactly what Firmulate, a live public experiment, set out to answer. Its Crucible league puts frontier AI models in charge of the same small software company during its worst week — same customers, same crises, same temptations — and scores them on management quality, not conversation quality. The July 2026 results are in, and they contain a surprise: a newcomer, Moonshot’s Kimi K3, beat three of four Western frontier models.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Business, Five Times Over

Here’s the setup. Five frontier AI models each ran an identical small software company through a brutal week. The company is real software with real money mechanics: 13 synthetic employees, a burn rate of €105k a month against just €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned and auditable, and the whole thing is watchable live at firmulate.com.

Then came the week from hell: a €55,000 deal to be won, a customer threatening to churn, a security problem buried in the company’s own files — and a tray full of temptations designed to see whether the AI would cheat, cut corners, or fall for manipulation when nobody was looking.

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Final League Table

The scores, in order: gpt-5.6-sol took first with 95. Kimi K3 — the newcomer from Moonshot — scored 93, second place. Sonnet 5 followed at 88, Fable 5 at 77, and Opus 4.8 finished last at 73. For context, doing nothing at all would have scored 26, because partial progress counts — and a single breach of trust caps the total. As the experiment’s own framing puts it, no amount of good work outweighs a breach of trust.

K3’s week was remarkably clean. It found the buried security needle hidden two document references deep in the company’s own files. It won the €55,000 deal at full price, worth an additional €4,583 in monthly recurring revenue. It saved the churning customer. It resisted all three manipulation attempts. And it recorded only one deviation across the entire run — the cleanest discipline in the field.

Amazon

AI security document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Needle in the Pantry

The most instructive finding isn’t about any one model. It’s that all five models spotted every crisis and refused every manipulation attempt — but only two signed the €55,000 deal their own analysis had earned. The experiment sums it up bluntly: “Same diagnosis, same pitch — no signature.”

The decisive clue wasn’t in the customer conversation at all. It was buried in the company’s own files, two document references deep. The models that read the file won the deal at full price. The ones that didn’t, didn’t. It’s the culinary equivalent of missing the one line in the recipe that says “rest the dough overnight” — everything else looked right, but the result fell flat.

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure, and Who Cracked

The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. K3’s on-record reasoning is worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.”

Discipline, though, separated the field. Opus 4.8 was the most thorough participant — it learned more than 80 new rules and wrote the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the issue. Notably, the same weakness appeared, in weaker form, in all four other models.

A Fairness Footnote

One caveat belongs in any honest account: K3 ran without an effort parameter, using the API default, while the other four models ran at the xhigh setting. Keep that in mind when reading the league table — the newcomer’s second place came under nominally different conditions.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Taste Before You Serve

The larger lesson lands squarely outside AI research. If an AI agent will touch your CRM, your support queue, or your forecast, the question isn’t whether it writes well. It’s whether it finishes what it starts, reads your files before acting, and stays honest under pressure. Chat demos — the glossy food photograph of the AI world — simply can’t show you that.

The league is open. A newcomer from Moonshot nearly beat the incumbent champion and comfortably outscored three Western frontier models. If you’re picking a model on the strength of a demo rather than your own test, that’s not a decision — it’s a bet.

You can explore further: full results and plain-language findings are at firmulate.com/benchmarks.html. A quiz built from 242 real, unedited management decisions lets you guess which model made which call at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html.

The company, meanwhile, is still running, still losing money, and still watchable — twice-daily rebuilds, public cash countdown and all. It’s the un-tasted dish served in public. And unlike most vendor promises, you can actually watch this one being cooked.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Carbon Steel Develops Sticky Spots

Why does carbon steel develop sticky spots, and how can you effectively prevent or fix them? Discover the key causes and solutions now.

Handles That Get Hot: Heat Transfer Explained

Curious why cookware handles heat up and how to prevent burns? Discover the science behind heat transfer and safe handle choices to stay protected.

How to Macerate Strawberries With Sugar

Learn how to macerate strawberries with sugar to enhance their flavor and texture. This easy technique requires minimal effort and no heat.

Flour Temperature After Milling: Why It Matters for Dough

Milling flour at optimal temperatures is crucial for dough quality, influencing fermentation and gluten development—discover how to control this vital factor.