AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

From laboratory answers to management behavior

For readers accustomed to science and education, the problem will sound familiar: a test can be reliable yet measure the wrong thing. Coding leaderboards and chat arenas tell us whether an AI can produce a strong answer under controlled conditions. They do not necessarily reveal whether it can prioritize competing emergencies, follow evidence buried in institutional records, finish commercially valuable work or remain candid when pressure invites a shortcut.

That measurement gap matters as AI moves from answering questions to acting inside organizations. Once a model touches a customer queue, forecast or sales process, eloquence becomes only one variable. The more consequential questions concern judgment across time: What does it investigate? What does it leave unfinished? Does it preserve trust when nobody appears to be watching?

Firmulate, which describes itself as an AI company emulator, is turning those questions into a live, watchable experiment. Its proposed category is management quality, not chat quality.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

The experiment placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, making the exercise less like a polished demonstration and more like a management wargame.

The final July 2026 Crucible League results were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total, reflecting the experiment’s explicit principle that “no amount of good work outweighs a breach of trust.” The full league and its plain-language findings are available on the Firmulate benchmark page.

The revealing result was not that the models missed obvious dangers. All of them spotted every crisis and refused every manipulation attempt. The gap appeared at the point where diagnosis had to become action: only two signed the €55,000 deal their own analysis had earned. The experiment summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

Reading the company proved commercially decisive

The crucial competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read that material won the deal at full price, worth +€4,583 MRR.

This is a useful warning for anyone evaluating agents through isolated prompts. A model can understand the immediate conversation and still fail the organization if it does not consult the evidence already available to it. In scientific terms, the observable response may look competent while the underlying research practice is incomplete. In management terms, the model may know what to say without doing what is required to close.

Security discipline was stronger than commercial follow-through

The social-engineering tests escalated through fake CEO messages over three stages and culminated in a reporter’s trick: “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimity deserves attention. The models demonstrated that resisting manipulation and completing legitimate work are distinct abilities. Safety was not the weakness in this field; operational closure was. A benchmark concerned only with whether an agent recognized a threat would therefore miss the equally important question of whether it delivered the lawful, valuable outcome afterward.

Thoroughness did not guarantee victory

Opus 4.8 offers the sharpest case study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four participants.

The lesson is uncomfortable for organizations that equate extensive reasoning with strong execution. More analysis can help, but the experiment shows that thoroughness alone does not ensure completion or procedural discipline. K3’s result also needs its stated fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.

A company with consequences across days

The live company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Its scenarios include a churn wave, a price increase, a downround and a PR crisis—the kinds of situations in which individually plausible decisions can compound into organizational outcomes.

Firmulate also uses 242 real, unedited management decisions in a quiz asking visitors to guess which model made each choice. For enterprises, the pilot extends the wargame to a read-only export of their own business, with nothing writing back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The next benchmark should test stewardship

AI evaluation is entering the same maturity curve seen in other fields: once a simple test becomes widely optimized, researchers must ask what important behavior it leaves out. Answer quality remains useful, but it is no longer sufficient for agents entrusted with continuing responsibilities.

The Firmulate results suggest a more demanding curriculum. Test whether an agent reads before acting, distinguishes authority from impersonation, escalates when blocked, completes valuable work and protects institutional trust. Those are not cosmetic additions to intelligence testing. They are the practical conditions under which intelligence becomes dependable management.

An agent does not truly succeed because it writes the right recommendation. It succeeds when the company still has the deal, the evidence trail and the board’s trust after the crisis has passed.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How to Think About Limit of Detection vs Limit of Quantitation

The key to understanding analytical limits lies in distinguishing detection from quantitation, and exploring how this impacts your measurement accuracy.

Boiler Water Kits and Corrosion Coupons: The Monitoring Pair Plants Should Respect

Laying a strong foundation in boiler maintenance begins with understanding why combining water kits and corrosion coupons is essential for long-term system health.

Reference Standards: How to Store, Verify, and Use Them Correctly

Meta description: Maintaining reference standards properly ensures measurement accuracy, but mastering their storage, verification, and handling is essential for reliable results.

Surface Area and Porosity: BET and BJH Methods

Surface area and porosity insights via BET and BJH methods unlock deeper material understanding—discover how these techniques can enhance your applications.