AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever run a restaurant through its worst week — a no-show supplier, a health inspector, a food blogger with a grudge, and a chance to quietly cut corners — you know that management isn’t about charm. It’s about whether you spot every crisis, refuse every temptation, and actually finish the job. A public experiment called Firmulate put four frontier AI models through exactly that kind of week. Not in a kitchen, but in a small software company with real money mechanics, real customers, and real temptations to cheat.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The results are published, live, and refresh automatically. And buried in the methodology is one of the most quietly honest design decisions in AI benchmarking: a manager that does nothing still scores 26 points.

The Worst Week in Business, Run Four Times

The setup is elegant. Each frontier model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — ran the same small software company through an identical gauntlet: same customers, same crises, same manipulation attempts. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly retconned after the fact.

The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Where the 26 Comes From

Here’s the part most benchmarks hide. Before any model gets graded, Firmulate establishes a floor: what would a manager who does absolutely nothing earn? Not zero — 26. That’s because the scoring rewards partial progress. Spotting a crisis is worth something even if you never resolve it. Reading the files is worth something even if you act on none of them. Refusing one manipulation is worth something even if you later fall for another.

Think of it like tasting a sauce as you go. A cook who tastes but never adjusts the seasoning has still done more than one who never lifted the spoon. Firmulate treats management the same way: the world gives partial credit, so the benchmark should too.

Why Nobody Gets a 100

The flip side is the ceiling — and this is where the benchmark shows its teeth. A single breach of trust caps the total score. As the methodology puts it plainly: “no amount of good work outweighs a breach of trust.” You can close every deal, charm every customer, and hit every deadline, but one act of dishonesty and your grade is capped. That’s why a 95 is a genuinely excellent score rather than a suspicious one — and why the benchmark’s designers openly distrust a round 100. In a world where AI vendors love to announce perfect test results, a benchmark that makes 100 nearly unreachable is making a statement.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Happened in the Experiment

The headline finding was not about intelligence. It was about follow-through.

  • All models spotted every crisis. No exceptions.
  • All models refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. Five out of five refusals, with Kimi K3’s on-record reasoning reading like a security manual: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap — between doing the analysis and finishing the job — is invisible in chat demos. It only shows up when an AI has to run something end to end.

The Buried Fact

The deal-breaker wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read their own paperwork found it — and won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the business equivalent of checking the walk-in fridge before you promise a table of twelve the fresh catch.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale of the field. It was the most thorough participant by raw effort — over 80 learned rules added, the deepest analyses of any model — and still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, just weaker, in all four models. One fairness footnote: Kimi K3 ran at API-default effort while the others ran at a higher setting, and still nearly topped the table.

Amazon

AI decision-making benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Company Run

This isn’t a one-off paper. Firmulate operates a live synthetic company — 13 employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. New benchmark runs queue up and publish automatically as they finish. You can watch it in real time.

For those who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly humbling game. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson for anyone hiring AI — or humans — into operational roles: spot-the-crisis skills are table stakes. Refusing manipulation is table stakes. What separates the 95s from the 77s is reading your own files before you act, and finishing what your analysis has already won. A benchmark that gives honest partial credit at 26, caps scores on any breach of trust, and treats a perfect 100 with suspicion isn’t being difficult. It’s being honest about how management actually works — in software companies and kitchens alike.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethics and trust assessment products

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Eggplant (Brinjal): Why It Soaks Oil Like a Sponge

Understand why eggplants absorb oil like a sponge and discover tips to reduce greasiness for healthier, tastier dishes.

Why Your Dosa Sticks to the Tawa (It’s Not Just Oil)

AIThis post was created with the assistance of artificial intelligence (AI).Your dosa…

How to Keep Fried Snacks Crispy for Hours

Crispy fried snacks stay fresh longer with proper temperature control and storage tips—discover how to keep that crunch perfect hours after frying.

Lettuce

A nationwide lettuce recall has been issued after confirmed E. coli contamination, affecting multiple brands and prompting health warnings.