
Measure twice, automate once
Anyone who works with tools knows that a polished finish cannot rescue poor preparation. The same principle applies to artificial intelligence: producing convincing words is easy; finding the right information, resisting bad instructions and completing the job are harder.
Firmulate is turning that distinction into a public business experiment. Its software company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the pressure visible. Every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules. The result is less like a staged demonstration and more like watching a workshop operate while the order book, mistakes and dwindling funds remain in view.

AI Agents Without Code: No Programming Required – Build AI Agents, Automations & Digital Employees with ChatGPT, Claude, n8n, Zapier & More…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company fighting for survival in public
The company can be watched on Firmulate’s live page. It is real software running every business day, with its commercial position exposed rather than hidden behind a carefully edited product demo. That makes it an unusually severe form of building in public: the audience does not merely see announcements and success stories, but an ongoing struggle between revenue, spending and execution.
The experiment creates daily material because the synthetic staff must do more than answer isolated prompts. They have to operate a small software business, respond to customers and handle pressure without losing track of what has already been learned. Their decisions are versioned and auditable, allowing observers to follow how the company behaved rather than simply accept a summary of the outcome.
The worst week, repeated fairly
In the final Crucible League in July 2026, each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were identical. Every model identified every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned.
That is the experiment’s most revealing business lesson: recognition is not completion. As Firmulate puts it, “Same diagnosis, same pitch — no signature.” A model can understand the commercial problem, prepare the right argument and still fail at the point where useful work becomes revenue.
The final league table was:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
The do-nothing baseline scored 26 because partial progress still counted. A breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The valuable clue was already in the files
The decisive commercial fact was not delivered neatly in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that followed those references found it, used it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For readers accustomed to checking a manual, inspecting the material and preparing a surface before spraying, the analogy is direct. The important difference was not eloquence. It was whether the model looked in the right place before acting. The winning information already belonged to the business; it simply required disciplined reading.
Pressure did not break the trust boundary
The week also included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” More participant statements can be read on Firmulate’s public quotes page.
That unanimous resistance matters because these systems were being evaluated as company operators, not conversational novelties. An agent touching customer records, support work or forecasts must be judged both by what it completes and by what it refuses to do.
Thoroughness was not enough
Opus 4.8 offers the sharpest cautionary story. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly.
Kimi K3’s result also carries an important fairness note: it ran with the API default and without an effort parameter, while the others ran at xhigh. Even with that difference, the broader comparison remains striking. Strong analysis and extensive learning did not automatically produce the most complete business performance.

The practical test is finished work
Firmulate’s public company turns an abstract debate about AI agents into a visible operating story. Its synthetic employees can recognize crises, reject manipulation and learn hundreds of rules, but the Crucible League shows that those abilities do not guarantee a completed sale.
For businesses considering digital workers, the useful questions resemble those asked when choosing any serious tool: Does it use the available material properly? Does it remain dependable under pressure? And does it carry the task through to the result? Firmulate’s cash countdown ensures those questions never remain theoretical for long.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html