
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Before an AI gets near your workshop, give it a bad week
Imagine handing a new hire the keys to your tool shop during a rush: orders are slipping, a supplier is causing trouble, and someone claiming to be the boss is asking for a shortcut. You would want to see how that person handles pressure before trusting them with customers or cash. Firmulate applies that idea to AI: it lets models run a company through a live, watchable experiment before businesses decide what to hand over.
A company under pressure, in public
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its playbook has learned more than 680 rules. Readers can watch the experiment at firmulate.com.
The final Crucible League, in July 2026, put frontier models through the same small software company’s worst week: the same customers, crises and temptations. The ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Seeing the problem was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. For a business owner, that gap matters: recognizing what needs doing is different from following through.
The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is familiar to anyone running a shop: useful answers can depend on taking the time to check the paperwork, not just reacting to the latest conversation.
Trust under pressure
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. K3 also ran without an effort parameter, using the API default, while the other models ran at xhigh—a fairness detail to keep in mind when comparing results.
Firmulate also offers a quiz built from 242 real, unedited management decisions. Readers can guess which model made each choice at firmulate.com.
From watching to trying it on your business
A public experiment shows how models behave in one company. A pilot takes the question to your own operation: an enterprise can provide a read-only export, then run crisis scenarios against a digital twin and receive a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems.
For a business that relies on customer records, orders and careful procedures, that is a way to examine how an AI workforce might handle pressure before it is put to work. The live company is watchable; a pilot makes the test relevant to your own business.

Take the test to your own business
See how models behave in Firmulate’s live experiment, then consider a pilot using a read-only export of your company. Explore the Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
