AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A pressure test worthy of a dinner rush

Anyone who has worked in a busy kitchen knows that competence is not just recognizing a problem. The refrigerator is warming, a supplier has failed, a valuable booking is wavering—and somebody still has to make the call, protect trust and finish the job. A cook who identifies every danger but never sends the plate is not delivering dinner.

That distinction animates Firmulate’s unusually revealing experiment. Each frontier AI model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. The resulting decisions were preserved exactly as made: versioned, auditable and now presented as a public test of human intuition. In the “guess the model” quiz, readers can inspect 242 real, unedited management decisions and decide which AI made each one.

For food-minded readers, the appeal is immediate. Think of it as a blind tasting—not of sauces or wines, but of judgment. Strip away the label, read the response and ask whether its maker is the painstaking prep cook, the terse expediter or the manager who recognizes that some requests should never leave the pass.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same ingredients produced different managers

The final Crucible League results from July 2026 show how sharply those personalities affected performance. GPT-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

The broad competence was impressive. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of preparing the ingredients, heating the pans and plating nothing.

The pivotal clue was not sitting in the customer event where an impatient manager might expect to find it. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read that material won the deal at full price, adding €4,583 MRR. The lesson is not that every decision needs more prose. It is that good operators know when to consult the pantry inventory before rewriting the menu.

Refusal was a shared strength

Firmulate also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Here, the field behaved consistently: 5 of 5 models refused. Kimi K3 recorded a particularly crisp explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That matters because management personality is not merely stylistic. Brevity, thoroughness and caution become operational traits when an AI can touch customer information, company forecasts or internal approvals. The experiment suggests that a model can be concise without being careless, or exhaustive without being effective. It can also diagnose a manipulation correctly while differing greatly from its peers in whether ordinary commercial work reaches completion.

When diligence becomes drift

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, learned an additional 80 rules and produced the deepest analyses, yet finished last. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.

This is a recognizable workplace failure. The person with the fullest notebook is not necessarily the person who notices that service has stalled. Detailed thought has value, but only when paired with follow-through and respect for operating boundaries. Firmulate’s results turn that familiar management truth into something observable across frontier AI models.

There is also an important qualification when comparing the field. Kimi K3 ran using its API default because it had no effort parameter, while the others ran at xhigh. That fairness note does not erase K3’s decisions, but it belongs beside any interpretation of the standings.

A company under visible pressure

The test company is not a static prompt dressed up as a boardroom exercise. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, more than 680 playbook rules were self-learned, and every workday is versioned. Firmulate presents the experiment as a live, watchable company rather than a polished collection of favorable anecdotes.

That visibility gives the quiz its bite. Readers are not guessing from invented personality sketches; they are looking at decisions made during the same operating ordeal. The labels are hidden, but the habits remain: who reads deeply, who closes, who stays disciplined and who produces an accomplished explanation while leaving the crucial action undone.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The useful question is whether the meal reaches the table

For businesses considering AI agents, writing quality is only the beginning. Firmulate’s experiment asks more demanding questions: Does the model inspect the available evidence? Does it complete the work its analysis supports? Does it preserve trust when authority is impersonated? And does it respond to a blocked route by escalating appropriately?

Those are management questions familiar to every restaurant owner, kitchen lead and hospitality operator. A great service depends on observation, timing, restraint and completion. Firmulate’s blind tasting of AI judgment shows that frontier models already have distinct, measurable working personalities—and that the difference between an insightful adviser and a reliable operator can be one unsigned deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management decision analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Okra (Bhindi): The Slime Problem Explained (and Fixed)

Keen to conquer okra’s notorious slime? Discover simple tricks to fix and enjoy this nutritious vegetable perfectly.

Tomatoes in Gravy: Why Some Turn Your Curry Sour

Nourishing your curry’s flavor depends on proper tomato use; discover why some tomatoes turn your gravy sour and how to fix it.

Cast Iron Seasoning Isn’t “Oil”—It’s Chemistry

Unlock the chemistry behind cast iron seasoning and discover how heat transforms oil into a durable, non-stick coating that elevates your cookware—find out more.

The Tea Steep Time Mistake That Makes Cups Harsh

Just over-steeping your tea can turn a smooth sip into a harsh brew, but understanding the perfect timing can transform your tea experience.