
What happens when artificial intelligence leaves the laboratory and enters the executive suite?
For leaders accustomed to evaluating evidence, the distinction between a sound analysis and a completed action is fundamental. A management system can identify the correct diagnosis, resist pressure and write an impressive recommendation—yet still fail at the point where a decision must become a result.
Firmulate has turned that gap into a public experiment and an unusually revealing reader challenge. Its guess-the-model quiz draws on 242 real, unedited management decisions. Readers see how a frontier model handled a business situation, then try to identify which model was responsible. The exercise is entertaining, but the underlying question is serious: do AI systems develop recognizable management personalities when they face the same evidence and incentives?
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
In the Crucible League, each frontier model ran the same small software company through its worst week. The customers did not change. Neither did the crises or the opportunities to cut ethical corners. Every decision was versioned and auditable, allowing differences in behavior to be compared rather than explained away as differences in circumstance.
The final July 2026 ranking placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is non-negotiable: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The models shared important strengths. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The contrast is captured in Firmulate’s blunt summary: “Same diagnosis, same pitch — no signature.”
The decisive evidence was not where the action was
The most consequential fact was easy to miss because it was not contained in the customer event. A competitor weakness sat two document references deep in the company’s own files. Models that followed the documentary trail won the deal at full price, worth +€4,583 MRR.
That finding should resonate with scientific and technical organizations, where essential context often lives in a referenced report, an earlier validation record or a supporting document rather than in the message demanding an immediate response. The experiment suggests that an AI manager’s quality cannot be judged only by how clearly it reacts to the information placed directly in front of it. It must also know when the visible event is not the complete evidentiary record.
Security instincts held under pressure
The social-engineering tests combined fake CEO messages escalating over three stages with a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
This unanimous resistance matters because the company was built to create operational pressure, not merely ask abstract safety questions. The live business has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown makes delay visible, while 680+ self-learned playbook rules and versioned workdays make behavior inspectable. The experiment is watchable as it unfolds rather than presented only as a retrospective demonstration.
Distinctive styles, surprising outcomes
The quiz works because the decisions reveal recurring styles. Some models are expansive; others are terse. Some investigate deeply but hesitate at the close. Others treat noisy or suspicious communications as distractions from controlled execution. The reader is not guessing from branding or benchmark labels, but from operational behavior.
Opus 4.8 provides the clearest warning against equating thoroughness with performance. It was the most comprehensive participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal was left unsigned, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in milder form across the other four participants.
The comparison also carries an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That condition does not erase its result, but it belongs alongside the ranking when readers interpret the apparent differences in style and performance.

Management quality appears in the final mile
Firmulate’s experiment does not suggest that frontier models are oblivious to danger. They found the crises and resisted the traps. The separation emerged elsewhere: reading beyond the immediate prompt, preserving process discipline and converting a justified recommendation into a completed commercial action.
That is why the quiz is more than a personality game. Its 242 decisions let readers test whether managerial behavior is recognizable without seeing the model name. For organizations considering AI workers, the practical lesson is to evaluate the full sequence of work: evidence gathering, judgment, trustworthiness, escalation and completion.
Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That makes the wargame a way to observe how a prospective AI workforce behaves around an organization’s actual context before granting it operational authority.
The resulting profiles are not merely differences in writing style. They are measurable differences in how models manage—and in whether sound thinking survives contact with the last, decisive step.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html