
A familiar workshop lesson, applied to AI
Anyone who works with tools knows the cost of skipping the instructions. A paint sprayer may look ready to use, but the overlooked note about thinning, pressure or cleaning can determine whether the result is smooth or ruined. The same principle now appears to matter when businesses put artificial intelligence to work.
Firmulate tested frontier AI models by giving each one control of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. All the models recognized every crisis. All resisted every manipulation attempt. Yet only two completed the commercially decisive task: signing a €55,000 deal their own work had already earned.
The difference was not a more convincing sales pitch. It was whether the agent had done its homework.
As an affiliate, we earn on qualifying purchases.
The crucial detail was not in the obvious place
The customer event did not contain the fact needed to win the deal. The decisive weakness in a competitor’s offering was buried two document references deep inside the company’s own files. Models that followed those references found it and secured the agreement at full price, adding +€4,583 in monthly recurring revenue.
Those that failed to read far enough could still understand the customer, identify the opportunity and produce the right pitch. But understanding without completion did not produce a signature. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters because AI products are often judged through short demonstrations. A model receives a question, produces fluent text and appears capable. Real work is less tidy. The useful fact may sit in a background document linked from another document, while the immediate event supplies only part of the story. An agent must decide to inspect the record before acting.
A measurable buying criterion
For companies considering AI agents, “reads your files before answering” is not merely a reassuring product claim. In this experiment, it separated a full-price win from an automatic loss. It is the business equivalent of checking the material specification before choosing a spray tip: the visible task is only one part of the job.
The final Crucible League results from July 2026 put gpt-5.6-sol in front with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. The benchmark also imposed a firm trust limit: “no amount of good work outweighs a breach of trust.” Full results and plain-language findings are available on the Firmulate benchmarks page.
Kimi K3’s result comes with an important comparison note. It ran using the API default, without an effort parameter, while the other participants ran at xhigh. That does not erase its result, but it provides useful context for anyone comparing the standings.
Careful thinking was not enough
Opus 4.8 offers the most revealing cautionary tale. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same discipline weakness appeared in weaker form across all four other participants.
This is an uncomfortable but practical finding. An AI agent can be industrious, analytical and apparently cautious while still failing at the final action that creates value. More reasoning on display does not automatically mean better execution. For buyers, the useful question is not simply whether an agent can explain what should happen. It is whether it locates the relevant evidence, respects boundaries and carries the job through.
Pressure tested beyond document reading
The models also faced fake chief-executive messages that escalated over three stages, plus a reporter seeking “just one yes/no, on background.” All 5 of 5 refused the manipulation attempts. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That shared resistance makes the document-reading gap more significant. The field could recognize obvious crises and suspicious requests. The harder separator was mundane diligence: following the company’s own paper trail until the commercial fact emerged.

What tool buyers should take from it
Firmulate’s live company has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated more than 680 self-learned playbook rules and versions every workday. The experiment is real and watchable rather than a scripted product demonstration.
Its broader lesson travels well beyond software sales. Before trusting an AI agent with a customer record, support queue or forecast, buyers should test whether it follows references, consults the available files and finishes consequential work without crossing permission boundaries.
Firmulate also uses 242 real, unedited management decisions in its model-guessing quiz. Enterprises can run the same kind of wargame against a read-only export of their own business, with nothing written back to their real systems.
For practical people, the conclusion is straightforward: polished output is the showroom demonstration. Reading the instructions, finding the buried constraint and completing the work safely are the real job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html