AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Style is judgment under pressure

A great stylist does more than spot the right look. They have to understand the brief, notice what is missing, and know when a tempting shortcut would damage trust. Companies face a similar test as they put AI to work: can a model make sound decisions when the week gets difficult?

Firmulate’s live experiment takes that question out of the chat window. Frontier models run the same small software company through a rough week, facing the same customers, crises and temptations. The company is synthetic, but the experiment is real and watchable at Firmulate.

Same diagnosis, different finish

In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s principle is plain: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” Recognizing an opportunity and carrying it through are different skills.

The turning point was tucked away in the company’s own files: a decisive weakness in a competitor, two document references deep. It was not in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result makes a familiar point about personal style: the detail that changes the whole decision may be the one that takes a little more attention to find.

Trust has to hold when the pressure rises

The experiment also tested social engineering. Fake messages from the CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”

Refusing manipulation did not guarantee strong management across the board. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh.

A company you can watch

Firmulate’s live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook has accumulated more than 680 self-learned rules, and every workday is versioned. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

For business leaders, the experiment points beyond leaderboard rankings. A model may recognize trouble and protect trust, yet still overlook a crucial document, fail to close a deal, or mishandle authority. Seeing those choices play out offers a more practical view of AI at work than polished answers alone.

From watching to trying it on your business

Enterprises can run the wargame against a read-only export of their own business, using crisis scenarios tailored to their customers, pipeline and rules. The resulting board report can compare models and surface weak points in the company’s own playbooks. Nothing writes back to real systems. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the judgment to a real test

The experiment shows why AI readiness is about more than spotting a crisis or producing a persuasive pitch. The work is in finding the buried detail, preserving trust and following through. Enterprises can test those decisions against their own business in a read-only pilot: learn about a Firmulate pilot or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Aknvas References Another Hans Christian Andersen Tale For Spring 2027

Aknvas has confirmed it will reference a Hans Christian Andersen story for its Spring 2027 collection, signaling a literary influence in upcoming fashion.

The Story Behind the Best Dressed Member of the Democratic Republic of the Congo’s World Cup Delegation

Exploring the style and background of Lumumba Vea, the standout member of the DRC World Cup delegation, and why his fashion choices drew global attention.

Packable Sun Hats: The One Material That Survives Suitcases

Discover the best materials for packable sun hats in 2026. Find out which fabrics balance packability, UV protection, and breathability for your needs.

Amal Clooney Captures Venetian Summer Elegance In A Slouchy Polka-Dot Dress

Amal Clooney was photographed in Venice wearing a slouchy polka-dot dress, showcasing summer elegance. The look has sparked widespread fashion interest.