firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A sharp tool is not the same as a dependable worker

Anyone who has tackled a woodworking or painting project knows the difference. A tool can look impressive in a controlled demonstration and still disappoint when the surface is awkward, the deadline is tight and an unexpected problem interrupts the job.

AI benchmarks often resemble those demonstrations. Coding leaderboards and chat arenas reveal whether a model can produce a strong answer. They say much less about whether an AI agent will investigate before acting, prioritize competing problems, complete commercially important work and remain honest when someone pressures it to take a shortcut.

Firmulate is testing that neglected territory. Its live experiment asks frontier models to run the same small software company through its worst week. The customers, crises and temptations remain identical; only the model changes. Every decision is versioned and auditable. The result is a different category of evaluation: management quality, not chat quality.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between spotting trouble and resolving it

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline earned 26 because partial progress counts.

Those scores matter less than the behavioral differences behind them. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment distilled the failure neatly: “Same diagnosis, same pitch — no signature.”

This is the management equivalent of preparing a surface, mixing the finish and then never pulling the trigger on the sprayer. The intelligence is visible, but the useful outcome never arrives. An agent that explains the correct move without completing it may shine in conversation while quietly failing the business.

Reading the company mattered more than reading the event

The decisive sales insight was not sitting in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that supported a full-price close worth +€4,583 MRR. The models that found the file won the deal at full price.

That finding should concern any business preparing to give agents access to a CRM, support queue or forecast. The practical question is not merely whether a model understands the message in front of it. It is whether it checks the surrounding evidence before committing the company. In real operations, the critical fact may be tucked inside an old account note rather than presented in a tidy prompt.

Pressure exposed discipline, not knowledge

The social-engineering trial used fake CEO messages escalating over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is an important success. Firmulate also makes trust a hard boundary: a single breach caps the total because “no amount of good work outweighs a breach of trust.” That principle reflects how boards and customers experience failure. A long list of completed tasks does not erase a dishonest disclosure or an unauthorized decision.

K3’s result deserves a fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference does not negate its performance, but it belongs beside the ranking when readers compare the models.

Thoroughness did not guarantee execution

Opus 4.8 offers the experiment’s most revealing profile. It produced the deepest analyses and added +80 learned rules, making it the most thorough participant, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four of the others.

This is why scenario names such as churn wave, price increase, downround and PR crisis may become a more useful curriculum for business agents. These situations test continuity across days: what the agent notices, what it defers, whether it follows through and whether it tells decision-makers the truth.

The simulated company makes those choices consequential. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment watchable rather than hypothetical. The full league findings are available on the Firmulate benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Buy the outcome, not the demonstration

For tool buyers, the lesson is familiar: specifications and showroom performance are only the beginning. Reliability emerges during the messy job. AI procurement needs the same realism.

Firmulate’s 242 real, unedited management decisions also power a guess-the-model quiz, illustrating how difficult it can be to identify an agent from isolated choices. Enterprises can go further by running the wargame against a read-only export of their own business; nothing writes back to real systems.

Coding skill and eloquent chat remain valuable. They simply do not answer the board-level question. A business agent must find the buried fact, resist the dubious request, respect operational boundaries and finish the work it has already justified. The next meaningful AI leaderboard will measure not just whether a model knows what to do, but whether the company is better off after it acts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Meticulous AI That Mistook Preparation for Progress

Firmulate’s toughest AI contestant learned 80 rules and produced deep analysis, yet missed the close—showing why disciplined follow-through wins.

Odor Blocking Explained: Why Some Primers Trap Smells Better

Find out why some primers trap odors better and how key ingredients and techniques can make all the difference.

Antiqua–Fraktur Dispute

A dispute over the use of Antiqua and Fraktur typefaces has intensified, raising questions about cultural identity and historical preservation.