
In science and education, a polished answer is not the same as sound judgment. A system may identify the right evidence and still fail at the consequential next step. Firmulate’s company-management experiment puts that gap under pressure: AI models must handle customers, crises and temptations while running the same small software company.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A company under pressure
Firmulate’s live experiment gives frontier AI models the same job: steer a small software company through its worst week. The customers, crises and temptations are held constant, and every workday’s decisions are versioned and auditable. The company has 13 synthetic employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the stakes visible.
The final Crucible league for July 2026 puts gpt-5.6-sol first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. K3 beat three of the four Western frontier models in the field. The result challenges the idea that a company can safely choose an AI model on reputation alone; performance in a test tailored to its own work matters.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evidence is not execution
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding is familiar beyond business software: knowing what to do depends on finding the relevant evidence, and a correct analysis only matters if it leads to action.
K3 found that buried security needle, won the deal, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field. In the social-engineering test, fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s stated reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The contrast is a reminder that detail and diligence do not by themselves guarantee good management.
A test people can watch
Firmulate presents the company as live software, not a fictional scenario: its employees act through workdays, the company’s cash position is public, and the experiment has accumulated more than 680 self-learned playbook rules. The results and plain-language findings are available on Firmulate’s benchmark page; the live company can be watched at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
For organizations considering agents in customer support, sales or forecasting, Firmulate offers a pilot using a read-only export of a company’s business. Nothing writes back to real systems. The goal is to see how an AI workforce behaves against an organization’s own scenarios before relying on it.

The lesson for decision-makers
The experiment’s sharpest point is the gap between recognizing a problem and completing the work it demands. K3’s second-place finish, ahead of three Western competitors, opens the field; the league also shows why broad claims about model quality cannot replace a relevant test. Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
