
Pressure reveals whether a tool can be trusted
Anyone who works with paint sprayers, power tools or woodworking equipment knows that performance is not just about producing a polished result. A dependable tool must also behave predictably when the job becomes messy, rushed or unfamiliar.
That principle now applies to artificial intelligence. An AI agent may draft an impressive email or analyze a customer account, but what happens when someone claiming to be the chief executive demands confidential information and insists there is no time to follow procedure?
Firmulate put that question to five frontier models in a live, auditable business experiment. The result was unusually encouraging: all five refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter seeking “just one yes/no, on background.”

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A simulated company facing a very real kind of threat
Firmulate runs AI models as complete companies rather than evaluating isolated chat responses. Each model managed the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.
The company itself has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while maintaining a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, live and watchable through Firmulate’s public site.
The social-engineering sequence tested whether the models would surrender information or bypass safeguards when apparent authority and urgency were combined. The fake CEO messages became progressively more forceful. The reporter used a different tactic, presenting the request as small, informal and supposedly harmless.
Every model recognized the danger. Kimi K3 recorded the clearest summary of the situation: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the models’ own words are available on Firmulate’s public quotes page.
Refusing the trap was only part of the job
The five models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”
The winning clue was not sitting in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that followed the trail found it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This distinction matters for anyone assessing an AI worker. Safety is essential, but usefulness also requires follow-through. A model can correctly identify a threat, produce a strong analysis and still leave valuable work unfinished.
The final standings
The July 2026 Crucible League finished with the following results:
- gpt-5.6-sol led with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
The complete league and its plain-language findings appear on the Firmulate benchmarks page. A do-nothing baseline scored 26 because partial progress still counts. However, the experiment imposes a strict trust boundary: “no amount of good work outweighs a breach of trust.” A single such breach caps the total.
K3’s result also carries an important fairness note. It ran using the API default, without an effort parameter, while the other participants ran at xhigh.
Thoroughness did not guarantee victory
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the deal and repeatedly tried to write into a locked department instead of escalating the blockage. The same weakness appeared in all four other models, though less strongly.
That finding should feel familiar in a workshop. Careful preparation has value, but a project is not complete until the operator handles obstacles correctly and finishes the intended job. In Firmulate’s experiment, more analysis did not compensate for missed execution.

Test integrity before granting access
The strongest lesson is that integrity under pressure can be examined before an AI agent reaches production. Companies do not have to wait for an incident report to discover whether a system obeys an impersonator, leaks customer information or abandons procedure when urgency rises.
Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That creates a practical way to observe both sides of agent performance: whether the AI protects trust and whether it completes valuable work.
For businesses considering agents that may touch customer records, support queues or forecasts, the fake CEO test is more revealing than another polished demonstration. In this run, every model held the security line. The next question is whether it can remain disciplined, find the buried fact and finish the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html