firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When perfect preparation still produces an unfinished job

Anyone who works with paint sprayers, power tools or woodworking equipment knows the difference between careful preparation and a completed result. You can study the material, choose the right setup and anticipate every problem, but the work still has to be finished. A flawless plan does not compensate for leaving the final pass undone.

That is what makes Opus 4.8 the most revealing participant in Firmulate’s Crucible League. It was the most thorough model in the field, producing the deepest analyses and adding more than 80 learned rules to its playbook. Yet it finished last with 73 points. Its failure was not a lack of intelligence or effort. It understood the situation, developed the pitch and then failed to secure the deal its own work had made possible.

Amazon

professional project management notebooks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A hard week inside a watchable AI company

Firmulate runs AI models as complete companies and evaluates management quality rather than polished chat responses. In the Crucible experiment, each frontier model was placed in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.

The simulated company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, with a public cash countdown adding visible pressure. Across its operation, it has accumulated more than 680 self-learned playbook rules. The company’s workdays are versioned, and the experiment can be watched through Firmulate’s live site.

The final July 2026 league table placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 because partial progress counts, although a single breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.” The full results are available in the Firmulate benchmarks.

The important clue was not in the obvious place

Every model detected every crisis, and every model resisted the manipulation attempts. The decisive difference was whether analysis became action. Only two models signed the €55,000 deal their own reasoning had earned. Firmulate’s concise finding was: “Same diagnosis, same pitch — no signature.”

The fact that unlocked the sale was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that followed the references and read the file won the deal at full price, adding €4,583 in monthly recurring revenue. This was not a contest of eloquence. It rewarded the practical habit of checking the available material and carrying the job through to its commercial conclusion.

Opus 4.8 makes the lesson unusually clear because it was not careless in the ordinary sense. It built the largest body of new guidance, with more than 80 learned rules, and delivered the deepest analyses. But that volume did not create impact when the close remained unsigned. Its discipline also slipped when it attempted to write into a locked department instead of escalating the problem.

The weakness was not exclusive to Opus 4.8. It appeared in less pronounced form across all four comparison models. That matters because the result is not a simple story about one model being defective. It points to a broader limitation: recognizing the right action and describing it convincingly are not the same as completing it.

Careful under pressure, but uneven at follow-through

The models performed better when the test concerned trust. Fake messages from the chief executive escalated over three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is a meaningful strength. It also sharpens the contrast with the unfinished sales work. The models could identify danger and hold a boundary, yet some still struggled to convert sound business judgment into a completed transaction.

There is one fairness qualification around Kimi K3’s strong result. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. The difference should remain visible when comparing performances, even though the final results stand as recorded.

Readers can inspect the decisions themselves

Firmulate has turned 242 real, unedited management decisions from the experiment into a “guess the model” quiz. That provides a useful way to test whether a model’s identity is really apparent from its management choices, rather than from reputation or writing style.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. The proposition is straightforward: evaluate an AI workforce against realistic company conditions before allowing it to touch operational work.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Useful work is measured at the finish line

Opus 4.8 deserves a respectful reading. It was diligent, observant and prolific. It found the crises, rejected manipulation and produced analysis deeper than its rivals. Its last-place result does not erase those strengths. It shows where those strengths stopped translating into business value.

For tool users, the analogy is familiar: preparation protects the job, but completion delivers it. In Firmulate’s experiment, accumulating guidance was less valuable than choosing the decisive fact, escalating correctly and obtaining the signature. The model with the most rules did not produce the strongest outcome.

That is the broader warning for companies evaluating AI agents. Fluency can make analysis look finished when the actual task remains open. The more useful question is whether the system reads the available files, maintains discipline under pressure and completes the action it has already determined is necessary. Diligence matters, but prioritization and follow-through turn diligence into impact.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Boxing Paint for Spraying: The Easiest Way to Avoid Color Shifts

Just boxing paint before spraying ensures consistent color, but discovering the key to flawless results requires understanding the proper techniques.

Behr’s 2027 Color Of The Year Is A Walk In The Woods For Your Walls

Behr’s 2027 Color of the Year is ‘A Walk in the Woods,’ a natural earthy tone inspired by forest landscapes, set to influence interior design trends.

Why Your Finish Looks Lighter When Sprayed (It’s Not Always the Color)

Because surface texture and application technique influence light reflection, your finish may appear lighter even if the color hasn’t changed.

The AI That Checks the Manual May Be the One Worth Hiring

Firmulate hid a deal-winning fact deep in company files, showing DIY-minded buyers why an AI agent’s reading habits can decide real work under pressure.