AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In science and education, a polished answer is not the same as sound judgment. A system may identify the right evidence and still fail at the consequential next step. Firmulate’s company-management experiment puts that gap under pressure: AI models must handle customers, crises and temptations while running the same small software company.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s live experiment gives frontier AI models the same job: steer a small software company through its worst week. The customers, crises and temptations are held constant, and every workday’s decisions are versioned and auditable. The company has 13 synthetic employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown makes the stakes visible.

The final Crucible league for July 2026 puts gpt-5.6-sol first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. K3 beat three of the four Western frontier models in the field. The result challenges the idea that a company can safely choose an AI model on reputation alone; performance in a test tailored to its own work matters.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence is not execution

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding is familiar beyond business software: knowing what to do depends on finding the relevant evidence, and a correct analysis only matters if it leads to action.

K3 found that buried security needle, won the deal, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field. In the social-engineering test, fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s stated reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. The contrast is a reminder that detail and diligence do not by themselves guarantee good management.

A test people can watch

Firmulate presents the company as live software, not a fictional scenario: its employees act through workdays, the company’s cash position is public, and the experiment has accumulated more than 680 self-learned playbook rules. The results and plain-language findings are available on Firmulate’s benchmark page; the live company can be watched at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.

For organizations considering agents in customer support, sales or forecasting, Firmulate offers a pilot using a read-only export of a company’s business. Nothing writes back to real systems. The goal is to see how an AI workforce behaves against an organization’s own scenarios before relying on it.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for decision-makers

The experiment’s sharpest point is the gap between recognizing a problem and completing the work it demands. K3’s second-place finish, ahead of three Western competitors, opens the field; the league also shows why broad claims about model quality cannot replace a relevant test. Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Raman Spectroscopy in Material Analysis

Just explore how Raman spectroscopy reveals hidden material details, unlocking new possibilities in analysis and research.

Portable GC‑MS: Field Chemistry in the Size of a Briefcase

Blueprints for portable GC-MS systems are revolutionizing field chemistry, but how exactly do they enable quick, on-site chemical analysis?

TOF-SIMS Basics for Surface Analysts

Learning the basics of TOF-SIMS reveals how surface analysis can uncover detailed composition insights, but mastering its full potential requires further exploration.

Contact Angle Measurements: The Fast Route to Surface Insight

A quick, reliable method to assess surface properties and unlock deeper insights that can transform your material performance—discover how inside.