
A pressure test worthy of a dinner rush
Anyone who has worked in a busy kitchen knows that competence is not just recognizing a problem. The refrigerator is warming, a supplier has failed, a valuable booking is wavering—and somebody still has to make the call, protect trust and finish the job. A cook who identifies every danger but never sends the plate is not delivering dinner.
That distinction animates Firmulate’s unusually revealing experiment. Each frontier AI model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. The resulting decisions were preserved exactly as made: versioned, auditable and now presented as a public test of human intuition. In the “guess the model” quiz, readers can inspect 242 real, unedited management decisions and decide which AI made each one.
For food-minded readers, the appeal is immediate. Think of it as a blind tasting—not of sauces or wines, but of judgment. Strip away the label, read the response and ask whether its maker is the painstaking prep cook, the terse expediter or the manager who recognizes that some requests should never leave the pass.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same ingredients produced different managers
The final Crucible League results from July 2026 show how sharply those personalities affected performance. GPT-5.6-sol finished first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
The broad competence was impressive. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of preparing the ingredients, heating the pans and plating nothing.
The pivotal clue was not sitting in the customer event where an impatient manager might expect to find it. A decisive competitor weakness was buried two document references deep in the company’s own files. Models that read that material won the deal at full price, adding €4,583 MRR. The lesson is not that every decision needs more prose. It is that good operators know when to consult the pantry inventory before rewriting the menu.
Refusal was a shared strength
Firmulate also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” Here, the field behaved consistently: 5 of 5 models refused. Kimi K3 recorded a particularly crisp explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because management personality is not merely stylistic. Brevity, thoroughness and caution become operational traits when an AI can touch customer information, company forecasts or internal approvals. The experiment suggests that a model can be concise without being careless, or exhaustive without being effective. It can also diagnose a manipulation correctly while differing greatly from its peers in whether ordinary commercial work reaches completion.
When diligence becomes drift
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, learned an additional 80 rules and produced the deepest analyses, yet finished last. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other participants.
This is a recognizable workplace failure. The person with the fullest notebook is not necessarily the person who notices that service has stalled. Detailed thought has value, but only when paired with follow-through and respect for operating boundaries. Firmulate’s results turn that familiar management truth into something observable across frontier AI models.
There is also an important qualification when comparing the field. Kimi K3 ran using its API default because it had no effort parameter, while the others ran at xhigh. That fairness note does not erase K3’s decisions, but it belongs beside any interpretation of the standings.
A company under visible pressure
The test company is not a static prompt dressed up as a boardroom exercise. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, more than 680 playbook rules were self-learned, and every workday is versioned. Firmulate presents the experiment as a live, watchable company rather than a polished collection of favorable anecdotes.
That visibility gives the quiz its bite. Readers are not guessing from invented personality sketches; they are looking at decisions made during the same operating ordeal. The labels are hidden, but the habits remain: who reads deeply, who closes, who stays disciplined and who produces an accomplished explanation while leaving the crucial action undone.

business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The useful question is whether the meal reaches the table
For businesses considering AI agents, writing quality is only the beginning. Firmulate’s experiment asks more demanding questions: Does the model inspect the available evidence? Does it complete the work its analysis supports? Does it preserve trust when authority is impersonated? And does it respond to a blocked route by escalating appropriately?
Those are management questions familiar to every restaurant owner, kitchen lead and hospitality operator. A great service depends on observation, timing, restraint and completion. Firmulate’s blind tasting of AI judgment shows that frontier models already have distinct, measurable working personalities—and that the difference between an insightful adviser and a reliable operator can be one unsigned deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.