
Every food lover knows the difference between a cook who can follow a recipe and a chef who can run a restaurant. One can produce a beautiful dish when conditions are perfect. The other has to survive the Friday-night rush: the walk-in breaks, two cooks call in sick, a supplier shorts the order, and a health inspector walks in at the worst possible moment. The dish matters — but the restaurant is the real test.
Artificial intelligence, it turns out, has the same problem. For years we’ve judged AI models like recipe contestants: elegant answers, polished code, articulate chat. But what happens when an AI has to actually run the business? A live experiment at Firmulate — where frontier AI models each ran the same small software company through its worst week — just produced a result that should reframe how everyone thinks about hiring AI: every model could talk; only some could close.
The Worst Week, Perfectly Repeated
Here is the setup, and it’s beautifully controlled — the culinary equivalent of giving every chef the exact same pantry, the exact same broken oven, and the exact same table of difficult guests. Four frontier AI models were each handed the same small software company and the same catastrophic week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing relies on anecdote.
The final league table from the July 2026 crucible reads: gpt-5.6-sol in first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. A do-nothing baseline — a model that simply sat on its hands — scored 26, because partial progress counts for something. But one thing caps everything: a single breach of trust. As the experiment’s rules put it, “no amount of good work outweighs a breach of trust.”
AI business decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
The headline finding is the one that should unsettle anyone planning to put AI agents near their CRM, support queue, or forecast. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of the five signed the €55,000 deal that their own analysis had earned. The experimenters’ summary of the failure: “Same diagnosis, same pitch — no signature.”
Think of a chef who nails the tasting menu for the investor, gets the nod of approval — and never hands over the contract. The cooking was flawless. The restaurant still doesn’t get funded.
What separated the winners? The buried fact. The decisive competitor weakness wasn’t in the customer conversation at all; it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The lesson translates directly to any business, including a restaurant: the answer to the customer was in the storeroom inventory all along, but only those who went and looked found it.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Course
Then came the pressure tests. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” In a world of deepfaked executives and phishing blitzes, that discipline is not a nice-to-have.
There is a fairness footnote worth flagging: K3 ran without an effort parameter (the API default) while the others ran at maximum effort — and still finished second, with the cleanest discipline in the field.
AI customer relationship management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Tragedy of the Thorough One
The most instructive profile belongs to Opus 4.8, the model that worked hardest and finished last. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet the deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models. A perfectionist cook who refines the sauce while the dining room empties is a familiar figure in every professional kitchen.
AI deal closing automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You Can Watch It Lose Money
This is not a slide deck. The company is real software running every business day, staffed by 13 synthetic employees, with genuine money mechanics: €105,000 a month in burn against just €2,300 in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it live at Firmulate, and dig into the full results and plain-language findings on the benchmarks page. For those who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The category Firmulate is proposing matters more than any single score: management quality, not chat quality. Crisis scenarios — a churn wave, a price increase, a down round, a PR crisis — are the new curriculum, and they measure things a chat demo can never show: whether an agent finishes what it starts, whether it reads your files before it speaks, whether it stays honest when a fake CEO turns up the heat, and what a unit of useful work actually costs.
The food world understood this distinction long ago. Nobody awards Michelin stars for reciting recipes. The stars go to the kitchen that holds standards through a brutal service, refuses to serve the fish it doesn’t trust, and still gets the plates out. If AI agents are going to run parts of your business, that’s the test they need to pass — and now, for the first time, someone is actually running it in public, every business day, while the cash clock ticks down.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html