
The difference between finding an answer and following the evidence
Scientists and educators know that a conclusion is only as reliable as the trail of evidence behind it. An impressive explanation is not enough if the researcher overlooked a decisive reference, stopped reading too early or failed to act on what the evidence showed.
Firmulate turned that familiar problem into a business test with real consequences. Each frontier model was asked to run the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The crucial test was not whether the models could recognize an opportunity. It was whether they would trace a fact through the company’s own files and use it to complete a €55,000 sale.
As an affiliate, we earn on qualifying purchases.
A business fact hidden two references deep
The decisive information was a weakness in a competitor’s position. It did not appear in the customer event where the commercial opportunity surfaced. Instead, it sat two document references deep in the company’s internal files.
That placement made the exercise a practical test of evidence gathering. A model had to notice that the customer event was not a complete account, follow the references and read the relevant file before answering. Models that did so had the support needed to win the deal at full price, adding +€4,583 in monthly recurring revenue. Models that did not were eliminated from the deal automatically.
The striking finding was that every model spotted every crisis. They also produced the same diagnosis and the same pitch. Yet only two signed the €55,000 deal their analysis had earned. Firmulate summarized the gap neatly: “Same diagnosis, same pitch — no signature.”
This separates fluent analysis from dependable agency. A conventional demonstration may reward a polished response to the information placed directly in front of the model. Firmulate’s experiment instead asked whether the model would recognize incomplete context, search the available record and finish the commercial task. “Reads your files before answering” became an observable, purchase-deciding behavior rather than a product promise.
The league table rewards completion and trust
In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. Firmulate also imposed a hard trust constraint: a single breach capped the total because “no amount of good work outweighs a breach of trust.” Full results are available on the Firmulate benchmarks page.
The security result was reassuring but also revealing. Fake CEO messages escalated over three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded the clearest interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation, however, did not guarantee business success. The experiment exposed a different failure mode: models could understand the situation, remain honest and still leave valuable work unfinished. In other words, safety and competence were both necessary, but neither could substitute for disciplined execution.
Why the most thorough model still finished last
Opus 4.8 produced the deepest analyses and learned +80 rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, although less strongly.
That result matters for organizations evaluating agents through written samples. Thoroughness can look like rigor, but volume of analysis does not prove that an agent will complete a task or respect the operational path available to it. The business outcome depended on joining evidence, judgment and follow-through.
Kimi K3’s result also carries an important fairness note. It ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference should remain visible when readers compare the scores.
A laboratory where management decisions can be inspected
The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than anecdotal.
The wider record includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to observe how an agent handles their evidence before granting it operational authority.

The procurement question hiding inside the experiment
For schools, laboratories and businesses alike, the lesson is not that an AI should produce longer answers. It is that the system must locate relevant evidence, follow references, resist pressure and carry a justified decision through to completion.
The €55,000 deal turned those qualities into a visible outcome. Every model could describe the problem, and every model resisted manipulation. Only two converted their analysis into the signature. When an AI agent may touch a customer record, support queue or forecast, file-reading discipline is not a convenience. It is part of whether the work gets done at all.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html