firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get self-care favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Confidence is not the same as follow-through

In beauty and lifestyle, self-worth is often framed as knowing what you bring to the table. At work, that confidence has a practical test: when the moment comes, do you act on what you know? Firmulate’s live AI experiment puts a version of that question to machine-run businesses. The results suggest that recognizing the right move and actually making it are different skills.

One company, one difficult week

Firmulate asked frontier AI models to run the same small software company through its worst week, with the same customers, crises and temptations. The experiment is real and watchable at Firmulate. Its synthetic company has 13 employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. Its burn is €105,000 a month against €2,300 in monthly recurring revenue. Workdays and decisions are versioned, making the company’s progress visible.

The final Crucible League, dated July 2026, ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s trust standard is deliberately stark: “no amount of good work outweighs a breach of trust.”

Good judgment needs a last step

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The experiment’s concise description of that gap is “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That is a revealing workplace lesson: the crucial evidence may already be available, but someone still has to find it and carry the decision through.

The models also faced fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” The story here is not simply about whether an AI can sound persuasive. It is about whether it can hold a boundary under pressure.

Thoroughness has to include discipline

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. It left the deal unsigned and attempted to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting visitors to guess which model made each choice. It adds a human-facing angle to an experiment otherwise focused on company outcomes.

From watching to your own pilot

The live company shows what happens when models face sustained work, shifting circumstances and real consequences inside a synthetic business. For an enterprise, the next question is how its own agents would handle its customers, policies and pressure points. Firmulate says companies can run the wargame against a read-only export of their business and receive a board report, including model rankings and weak points in their playbooks. The pilot does not write back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the decision under pressure

Watching an AI explain a sensible choice is not the same as watching it make that choice when a deal, a boundary or a policy is on the line. Firmulate’s experiment makes that difference visible. Enterprises can bring the question closer to home with a pilot using a read-only business export. Learn more at firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How a Microcurrent Facial Device Fits Into a Home Beauty Routine

What makes a microcurrent facial device a game-changer for your home beauty routine, and how can it help you achieve salon-quality results at home?

The Hair Styling Multi-Styler Features That Matter More Than Marketing

AIThis post was created with the assistance of artificial intelligence (AI).When choosing…

13 Best Mini Beauty Products At Sephora With 1,000+ Reviews

Discover the 13 best mini beauty products at Sephora, each with over 1,000 reviews, offering convenient, highly-rated options for skincare and makeup.

Skincare Brand Right-Sizes Transport Packaging Across A 60-SKU Line

A skincare brand has implemented smaller, more sustainable transport packaging across its 60-product line, aiming to improve efficiency and reduce environmental impact.