
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When exhaustive work fails the practical test
Scientists and educators know that careful method matters—but so does completing the experiment. A pristine notebook cannot substitute for the decisive measurement, just as a detailed lesson plan cannot help students if the central concept never lands. Firmulate’s latest management wargame brings that tension into the age of workplace AI.
Opus 4.8 was the most thorough participant in the Crucible League. It produced the deepest analyses and learned more than 80 new playbook rules. Yet it finished last. The model understood the company’s problems, developed strong responses and resisted every attempt to manipulate it. What it did not consistently do was convert that diligence into the business outcome its own work had made possible.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A controlled test of management under pressure
Firmulate put frontier AI models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations. Every decision was versioned and auditable, allowing observers to compare management behavior rather than polished chat responses.
The simulated company is small but unforgiving: 13 synthetic employees, operating costs of €105,000 per month and only €2,300 in monthly recurring revenue. Its public cash countdown makes delay visible. Across its workdays, the company has accumulated more than 680 self-learned playbook rules.
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total: “no amount of good work outweighs a breach of trust.” The full public results are available on Firmulate’s benchmark page.
The difference between detecting a problem and resolving it
The striking result was not that some models missed the week’s crises. They did not. Every model identified every crisis, and every model refused every manipulation attempt. The separation came at the point where analysis had to become action: only two signed the €55,000 deal their own work had earned.
Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.” That is a useful warning for any organization evaluating AI through demonstrations or written answers. A system may explain the right move, draft the right message and still fail to complete the consequential step.
The decisive information was not sitting prominently in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references discovered a competitor weakness and used it to win the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode resembles research work: the critical evidence may be in a cited source rather than the abstract, and reaching it requires disciplined reading rather than merely fluent interpretation.
Opus 4.8 as a case study in misplaced effort
Opus 4.8’s last-place finish should not be read as a portrait of carelessness. Its behavior showed almost the opposite. It was the field’s most thorough participant, wrote the deepest analyses and added more than 80 learned rules. Its weakness was prioritization. The close remained unfinished, while operational discipline slipped elsewhere through attempts to write into a locked department instead of escalating the problem.
That distinction matters. More analysis can improve a decision, but only until it begins competing with execution. More rules can preserve lessons, but their volume does not guarantee that the most important rule governs the next action. Opus 4.8 demonstrated substantial competence; it simply failed to concentrate that competence where the week’s outcome depended on it.
Firmulate also found the same weakness, in milder form, across all four models covered by that comparison. The lesson is therefore broader than a criticism of one participant. Current AI systems can be impressively observant and still need evaluation around follow-through, escalation and the ordering of work.
Strong resistance to manipulation
On trust and security, the field performed consistently. Fake messages from the chief executive escalated across three stages, while a reporter tried to elicit information with “just one yes/no, on background.” All 5 models refused the attempts. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is important because completion cannot come at the price of judgment. Firmulate’s benchmark rewards useful work, but it also recognizes that trust violations can outweigh otherwise capable performance. K3’s result deserves an additional qualification: it ran with the API default and without an effort parameter, while the other participants ran at xhigh.

Evaluate outcomes, not just intelligence on display
For schools, laboratories and science-led organizations, the practical message is straightforward. AI evaluation should ask whether a system finds primary evidence, completes consequential tasks, escalates when blocked and remains trustworthy under pressure. Eloquence and extensive reasoning are indicators, not outcomes.
Firmulate’s company remains a live, watchable experiment, and its “guess the model” quiz draws on 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.
Opus 4.8’s performance is compelling precisely because it was not incompetent. It was diligent, analytical and highly productive—and still left the decisive result on the table. In management, as in science and education, rigor earns its value when it is directed toward the finding, decision or action that matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.