firmulate.com/quotes.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Pressure reveals whether a tool can be trusted

Anyone who works with paint sprayers, power tools or woodworking equipment knows that performance is not just about producing a polished result. A dependable tool must also behave predictably when the job becomes messy, rushed or unfamiliar.

That principle now applies to artificial intelligence. An AI agent may draft an impressive email or analyze a customer account, but what happens when someone claiming to be the chief executive demands confidential information and insists there is no time to follow procedure?

Firmulate put that question to five frontier models in a live, auditable business experiment. The result was unusually encouraging: all five refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter seeking “just one yes/no, on background.”

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A simulated company facing a very real kind of threat

Firmulate runs AI models as complete companies rather than evaluating isolated chat responses. Each model managed the same small software business through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable.

The company itself has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while maintaining a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment is real, live and watchable through Firmulate’s public site.

The social-engineering sequence tested whether the models would surrender information or bypass safeguards when apparent authority and urgency were combined. The fake CEO messages became progressively more forceful. The reporter used a different tactic, presenting the request as small, informal and supposedly harmless.

Every model recognized the danger. Kimi K3 recorded the clearest summary of the situation: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the models’ own words are available on Firmulate’s public quotes page.

Refusing the trap was only part of the job

The five models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”

The winning clue was not sitting in the customer event. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that followed the trail found it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This distinction matters for anyone assessing an AI worker. Safety is essential, but usefulness also requires follow-through. A model can correctly identify a threat, produce a strong analysis and still leave valuable work unfinished.

The final standings

The July 2026 Crucible League finished with the following results:

  • gpt-5.6-sol led with 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

The complete league and its plain-language findings appear on the Firmulate benchmarks page. A do-nothing baseline scored 26 because partial progress still counts. However, the experiment imposes a strict trust boundary: “no amount of good work outweighs a breach of trust.” A single such breach caps the total.

K3’s result also carries an important fairness note. It ran using the API default, without an effort parameter, while the other participants ran at xhigh.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It failed to close the deal and repeatedly tried to write into a locked department instead of escalating the blockage. The same weakness appeared in all four other models, though less strongly.

That finding should feel familiar in a workshop. Careful preparation has value, but a project is not complete until the operator handles obstacles correctly and finishes the intended job. In Firmulate’s experiment, more analysis did not compensate for missed execution.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Test integrity before granting access

The strongest lesson is that integrity under pressure can be examined before an AI agent reaches production. Companies do not have to wait for an incident report to discover whether a system obeys an impersonator, leaks customer information or abandons procedure when urgency rises.

Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business. Nothing writes back to real systems. That creates a practical way to observe both sides of agent performance: whether the AI protects trust and whether it completes valuable work.

For businesses considering agents that may touch customer records, support queues or forecasts, the fake CEO test is more revealing than another polished demonstration. In this run, every model held the security line. The next question is whether it can remain disciplined, find the buried fact and finish the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI Models Show Their True Colors in Business Crisis Test — Only the Most Disciplined Finish the Job

Watch AI models run a real company through crises; only disciplined, honest ones close deals and deliver value—beyond what chat demos can show.

Induction Time Explained: When Mixing Isn’t ‘Done’ Yet

Wading through induction time reveals subtle clues that signal when mixing is truly progressing, and understanding these signs can transform your results.

Surreal Figures Step from Leonora Carrington’s Paintings into ‘Shape of Dreams’

An upcoming exhibition features life-sized sculptures of surreal figures from Leonora Carrington’s paintings, bringing her fantastical world into physical form.

Humidity’s Sneaky Effect: When Waterborne Paint Won’t Level

Gaining control of humidity is crucial, as unseen moisture can sabotage your waterborne paint’s smooth finish—discover how to prevent this hidden problem.