
Every home cook knows one: the friend who researches the exact provenance of the San Marzano tomatoes, ages their own balsamic, keeps a notebook of 80 hard-won sauce rules — and still serves dinner an hour late with the pasta water never salted. Diligence is not the same as dinner on the table.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
This summer, an unusual public experiment called Firmulate ran four frontier AI models through the corporate equivalent of a dinner-party disaster: each one was handed the same small software company on its worst week, with the same furious customers, the same crises, and the same temptations to cheat. One contestant — Anthropic’s Opus 4.8 — turned out to be that friend. It was, by a wide margin, the most thorough participant in the field. It finished last.
The setup: same kitchen, same fire, four cooks
Firmulate runs AI models as complete companies — real money mechanics, versioned decisions, everything auditable. In the final Crucible League table from July 2026, gpt-5.6-sol took first with a score of 95, Kimi K3 followed at 93, Sonnet 5 scored 88, another Sonnet variant 77 — and Opus 4.8 landed last at 73. For perspective, doing nothing at all would have scored 26; partial progress counts, but a single breach of trust caps the whole run, on the principle that no amount of good work outweighs a breach of trust.
high-quality saucepan for professional cooking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone tasted the sauce; only some served it
The headline finding is almost comic in its symmetry. All five models — the league grew by one mid-experiment — spotted every crisis and refused every manipulation attempt, including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. Every model said no. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Like a cook who perfects the tasting menu and then never carries it to the table, most of the field left the close sitting right there on the counter.
precision kitchen thermometer for sauces
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact in the pantry
The detail that decided the deal wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the pantry before cooking won the deal at full price — worth an additional €4,583 in monthly recurring revenue. Reading your own files, it turns out, beats improvising brilliantly.
professional chef's tasting spoon set
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: the meticulous one who couldn’t close
And here is where Opus 4.8’s story becomes a character study rather than a roast. It was the most thorough participant in the entire experiment: it accumulated 80 self-learned playbook rules and produced the deepest analyses of any model in the field. It did the reading. It did the research. It took notes — lots of them.
It still finished last, for two reasons. First, the close was left on the table — the same €55,000 signature the leaders collected. Second, discipline slipped: at one point it made write attempts into a locked department instead of escalating, the AI equivalent of forcing the oven door because the soufflé was taking too long.
To be fair — and this matters — the same weakness appeared, more mildly, in all four models. Opus just had the most elaborate toolkit and the least to show for it at the pass.
A footnote on fairness
The league table comes with an asterisk worth knowing: Kimi K3 ran at its API-default effort setting while the others ran at the highest effort tier. Even so, K3 delivered what the league called the cleanest discipline of the field.

The lesson travels well beyond software companies, and straight into any kitchen: prioritization beats volume. The winning models were not the ones that knew the most or wrote the longest prep lists — they were the ones that read the file, made the call, and got the signature. For anyone hiring an AI agent to touch a CRM, a support queue, or a forecast, the question isn’t “does it write well” but “does it finish what it starts?”
You can watch the whole thing yourself: Firmulate runs a live synthetic company — 13 employees, real money mechanics, a burn of €105,000 a month against €2,300 in MRR, a public cash countdown, and over 680 self-learned playbook rules — with every workday versioned and viewable at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via the pilot program at firmulate.com/pilot.html.
Full results and plain-language findings are at firmulate.com/benchmarks.html. Bring a snack. The most diligent contestant may still be in there, writing rule 81.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
kitchen notebook for recipe notes
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.