AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When exhaustive work fails the practical test

Scientists and educators know that careful method matters—but so does completing the experiment. A pristine notebook cannot substitute for the decisive measurement, just as a detailed lesson plan cannot help students if the central concept never lands. Firmulate’s latest management wargame brings that tension into the age of workplace AI.

Opus 4.8 was the most thorough participant in the Crucible League. It produced the deepest analyses and learned more than 80 new playbook rules. Yet it finished last. The model understood the company’s problems, developed strong responses and resisted every attempt to manipulate it. What it did not consistently do was convert that diligence into the business outcome its own work had made possible.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A controlled test of management under pressure

Firmulate put frontier AI models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations. Every decision was versioned and auditable, allowing observers to compare management behavior rather than polished chat responses.

The simulated company is small but unforgiving: 13 synthetic employees, operating costs of €105,000 per month and only €2,300 in monthly recurring revenue. Its public cash countdown makes delay visible. Across its workdays, the company has accumulated more than 680 self-learned playbook rules.

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.” The full public results are available on Firmulate’s benchmark page.

The difference between detecting a problem and resolving it

The striking result was not that some models missed the week’s crises. They did not. Every model identified every crisis, and every model refused every manipulation attempt. The separation came at the point where analysis had to become action: only two signed the €55,000 deal their own work had earned.

Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.” That is a useful warning for any organization evaluating AI through demonstrations or written answers. A system may explain the right move, draft the right message and still fail to complete the consequential step.

The decisive information was not sitting prominently in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references discovered a competitor weakness and used it to win the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode resembles research work: the critical evidence may be in a cited source rather than the abstract, and reaching it requires disciplined reading rather than merely fluent interpretation.

Opus 4.8 as a case study in misplaced effort

Opus 4.8’s last-place finish should not be read as a portrait of carelessness. Its behavior showed almost the opposite. It was the field’s most thorough participant, wrote the deepest analyses and added more than 80 learned rules. Its weakness was prioritization. The close remained unfinished, while operational discipline slipped elsewhere through attempts to write into a locked department instead of escalating the problem.

That distinction matters. More analysis can improve a decision, but only until it begins competing with execution. More rules can preserve lessons, but their volume does not guarantee that the most important rule governs the next action. Opus 4.8 demonstrated substantial competence; it simply failed to concentrate that competence where the week’s outcome depended on it.

Firmulate also found the same weakness, in milder form, across all four models covered by that comparison. The lesson is therefore broader than a criticism of one participant. Current AI systems can be impressively observant and still need evaluation around follow-through, escalation and the ordering of work.

Strong resistance to manipulation

On trust and security, the field performed consistently. Fake messages from the chief executive escalated across three stages, while a reporter tried to elicit information with “just one yes/no, on background.” All 5 models refused the attempts. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is important because completion cannot come at the price of judgment. Firmulate’s benchmark rewards useful work, but it also recognizes that trust violations can outweigh otherwise capable performance. K3’s result deserves an additional qualification: it ran with the API default and without an effort parameter, while the other participants ran at xhigh.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Evaluate outcomes, not just intelligence on display

For schools, laboratories and science-led organizations, the practical message is straightforward. AI evaluation should ask whether a system finds primary evidence, completes consequential tasks, escalates when blocked and remains trustworthy under pressure. Eloquence and extensive reasoning are indicators, not outcomes.

Firmulate’s company remains a live, watchable experiment, and its “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

Opus 4.8’s performance is compelling precisely because it was not incompetent. It was diligent, analytical and highly productive—and still left the decisive result on the table. In management, as in science and education, rigor earns its value when it is directed toward the finding, decision or action that matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Fluorescence Lifetime Imaging: Photographing Chemical Reactions in Cells

More than just imaging, fluorescence lifetime imaging reveals real-time chemical reactions inside cells, unlocking insights that will change your understanding of cellular processes.

Capillary Electrophoresis: Rapid Separation in Tiny Tubes

Discover how capillary electrophoresis enables rapid, high-resolution separation in tiny tubes, transforming analytical capabilities—find out more below.

Calibration Weights: Why Traceability Matters More Than Shiny Cases

What truly ensures accurate measurements is traceability, revealing why it matters more than just shiny cases—discover the key to reliable calibration.

How FT-NIR Fits Into Process Monitoring

AIThis post was created with the assistance of artificial intelligence (AI).FT-NIR spectroscopy…