AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A kitchen can look calm while a dinner service is going off course: a supplier misses a delivery, a reservation changes, and someone asks for an exception. The same kind of pressure arrives in a small business, where one bad decision can affect customers, cash and trust. Firmulate puts AI models through that pressure in a live, watchable company experiment.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A rough week, held constant

In the final Crucible League, run in July 2026, each frontier model was given the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The experiment asked a practical question: could a model do more than identify trouble? Could it act on its own analysis, close a sale and maintain discipline?

The league placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The standard included a firm boundary: a single breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

Seeing the problem is not the same as solving it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The concise description of the gap: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a telling distinction for any operator: the answer may depend on connecting evidence already in the business, then following through.

The trust tests were equally direct. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness has to reach the finish line

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a snapshot of this experiment, with that difference part of the record.

From watching to trying it on your business

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and versioned workdays. The site also offers a quiz built from 242 real, unedited management decisions: guess which model made each call.

For enterprises, the next step is a pilot using a read-only export of their own business. The company’s customers, pipeline and rules can anchor crisis scenarios and a board report showing model rankings and weak points in existing playbooks. Nothing writes back to real systems. It is a way to rehearse before entrusting AI with operational decisions.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Make the rehearsal specific

A model that diagnoses a problem still has to act, respect boundaries and carry the work through. Firmulate makes those decisions visible in a live company experiment, then offers enterprises a way to run the exercise against their own business data in read-only form.

Explore a Firmulate pilot or contact contact@firmulate.com to discuss wargaming your company.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Mixing Step That Decides Cake Texture

AIThis post was created with the assistance of artificial intelligence (AI).The key…

Cooking Oil Smoke Point: What It Means (and What It Doesn’t)

Permitting you to cook at optimal temperatures, understanding the smoke point reveals essential insights—yet there’s more to discover about choosing the right oil.

Why Some Induction Cooktops “Pulse” Heat (and How to Cook Through It)

Many induction cooktops pulse heat for precise control—discover why this happens and how to cook smoothly through it.

The Real Reason Food Browns Instead of Just ‘Cooking’

Savor the science behind why food browns and how mastering the Maillard reaction can elevate your cooking to new flavors.