AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Before an AI gets the keys, give it a bad week

In science and chemical businesses, a routine disruption can quickly become a customer, cash-flow or reputation problem. The useful question about AI is not just whether it can explain a procedure. Can it make sound decisions when several things go wrong at once—and hold to company rules under pressure?

Firmulate puts that question into a live experiment. Its public company emulator lets visitors watch synthetic employees make decisions, with real-money mechanics and a public cash countdown. The experiment is real and watchable at firmulate.com.

One company, the same worst week

In the final Crucible League, dated July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The final ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The headline finding was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came at the moment of execution: only two signed a €55,000 deal that their own analysis had earned. “Same diagnosis, same pitch — no signature.” A polished recommendation, in other words, does not guarantee follow-through.

The clue was buried in the company’s own files

The decisive competitor weakness sat two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder for any organization with valuable knowledge spread across documents: useful judgment depends on finding and applying what the company already knows.

Integrity faced a different test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs discipline

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. For leaders evaluating AI, that distinction matters: analysis, authority and operational discipline are separate parts of performance.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The result is a useful experiment to examine, not a reason to ignore how test conditions shape comparisons.

From watching to your own company

The live company has 13 synthetic employees, burns €105k/month against €2.3k MRR, and publishes a cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Visitors can also test their instincts against 242 real, unedited management decisions in the “guess the model” quiz at firmulate.com.

For an enterprise, the next step is a pilot using a read-only export of its own business. Leaders can put crisis scenarios against a digital twin and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That makes the exercise a way to examine how an AI workforce might handle the pressures of a specific business before giving it access to live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the judgment to the test

Firmulate’s experiment suggests that spotting a crisis and refusing manipulation are not the whole job: models also need to find buried evidence, close the work their analysis supports and respect operational boundaries. Enterprises can run the wargame against a read-only export of their own business. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Contact Angle Measurements: The Fast Route to Surface Insight

A quick, reliable method to assess surface properties and unlock deeper insights that can transform your material performance—discover how inside.

Chromatography 101: Separating the Unsavory From the Essential

With chromatography techniques, learn how to effectively separate impurities from essential compounds and unlock the secrets to perfecting your purification process.

Moisture Analyzers: Why Halogen Results Don’t Match Oven Drying

Just understanding the differences between halogen moisture analyzers and oven drying is crucial for accurate results and consistent quality.

Acoustic Imaging Cameras Explained: Why They Find Leaks So Fast

Proven to quickly locate leaks by translating sound waves into precise images, acoustic imaging cameras reveal hidden issues—discover how they boost detection accuracy.