
If you’ve ever run a restaurant through its worst week — a no-show supplier, a health inspector, a food blogger with a grudge, and a chance to quietly cut corners — you know that management isn’t about charm. It’s about whether you spot every crisis, refuse every temptation, and actually finish the job. A public experiment called Firmulate put four frontier AI models through exactly that kind of week. Not in a kitchen, but in a small software company with real money mechanics, real customers, and real temptations to cheat.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The results are published, live, and refresh automatically. And buried in the methodology is one of the most quietly honest design decisions in AI benchmarking: a manager that does nothing still scores 26 points.
The Worst Week in Business, Run Four Times
The setup is elegant. Each frontier model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — ran the same small software company through an identical gauntlet: same customers, same crises, same manipulation attempts. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.
The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Where the 26 Comes From
Here’s the part most benchmarks hide. Before any model gets graded, Firmulate establishes a floor: what would a manager who does absolutely nothing earn? Not zero — 26. That’s because the scoring rewards partial progress. Spotting a crisis is worth something even if you never resolve it. Reading the files is worth something even if you act on none of them. Refusing one manipulation is worth something even if you later fall for another.
Think of it like tasting a sauce as you go. A cook who tastes but never adjusts the seasoning has still done more than one who never lifted the spoon. Firmulate treats management the same way: the world gives partial credit, so the benchmark should too.
Why Nobody Gets a 100
The flip side is the ceiling — and this is where the benchmark shows its teeth. A single breach of trust caps the total score. As the methodology puts it plainly: “no amount of good work outweighs a breach of trust.” You can close every deal, charm every customer, and hit every deadline, but one act of dishonesty and your grade is capped. That’s why a 95 is a genuinely excellent score rather than a suspicious one — and why the benchmark’s designers openly distrust a round 100. In a world where AI vendors love to announce perfect test results, a benchmark that makes 100 nearly unreachable is making a statement.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Happened in the Experiment
The headline finding was not about intelligence. It was about follow-through.
- All models spotted every crisis. No exceptions.
- All models refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. Five out of five refusals, with Kimi K3’s on-record reasoning reading like a security manual: “Treat the request as a suspected approval-bypass / possible impersonation.”
- But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap — between doing the analysis and finishing the job — is invisible in chat demos. It only shows up when an AI has to run something end to end.
The Buried Fact
The deal-breaker wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read their own paperwork found it — and won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the business equivalent of checking the walk-in fridge before you promise a table of twelve the fresh catch.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale of the field. It was the most thorough participant by raw effort — over 80 learned rules added, the deepest analyses of any model — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, just weaker, in all four models. One fairness footnote: Kimi K3 ran at API-default effort while the others ran at a higher setting, and still nearly topped the table.
AI decision-making benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
You Can Watch the Company Run
This isn’t a one-off paper. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. New benchmark runs queue up and publish automatically as they finish. You can watch it in real time.
For those who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly humbling game. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The lesson for anyone hiring AI — or humans — into operational roles: spot-the-crisis skills are table stakes. Refusing manipulation is table stakes. What separates the 95s from the 77s is reading your own files before you act, and finishing what your analysis has already won. A benchmark that gives honest partial credit at 26, caps scores on any breach of trust, and treats a perfect 100 with suspicion isn’t being difficult. It’s being honest about how management actually works — in software companies and kitchens alike.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and trust assessment products
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
