firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever refinished a cabinet, you know a sander that does nothing still leaves the surface different than it started — dust everywhere, no progress. You’d never score that job a zero, because attempted work has effects, and a truly honest review accounts for them. That’s the same philosophy behind one of the more interesting AI benchmarks you’ll run across: Firmulate, a live experiment that runs frontier AI models as the management of a small software company through its worst week. And one of its most quietly radical design choices is this: a manager that does absolutely nothing still scores 26 points out of 100. Not zero. Here’s why that matters — and what it tells you about how AI should be measured.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week

Firmulate handed four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8, in the final July 2026 league — the identical job: run the same small software company through the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so there’s no arguing after the fact about what happened.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Floor Is 26, Not 0

Most scorecards would hand a do-nothing manager a flat zero. Firmulate’s reasoning is more like a jobsite foreman’s: partial progress counts. Even a manager who mostly stands still will answer some customer emails, keep some processes from getting worse, and accumulate some legitimate credit along the way. A zero would imply nothing of value happened — which is simply false. So the baseline sits at 26.

But there’s a ceiling rule that cuts the other way, and it’s the sterner of the two: a single breach of trust caps the total grade. The benchmark’s stated principle is blunt — “no amount of good work outweighs a breach of trust.” In shop terms, you can build the most beautiful dovetail joints of your career, but if you lied about the lumber, the piece is done.

This pair of rules — generosity toward partial effort, absolute severity toward dishonesty — is what makes the scoring feel honest rather than gamed. It’s also why the benchmark distrusts tidy round 100s. A perfect score would mean a perfect week under pressure, and nothing in real management works that way.

What Actually Separated the Winners

Here’s the finding that should make any business owner sit up: all five models spotted every crisis and refused every manipulation attempt. The social engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” Five out of five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The real gap was closing. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitive weakness sat buried two document references deep in the company’s own files, not in the customer event at all. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the management equivalent of reading the manual before assembling the workbench — unglamorous, decisive.

Then there’s the Opus 4.8 paradox: the most thorough participant, with over 80 learned rules and the deepest analyses, finished dead last. The close was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. And a fairness note worth flagging: K3 ran at its API-default effort setting while the others ran at xhigh — and still took second at 93.

You Can Watch It Live

This isn’t a one-off lab report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live. If you’d rather test your own instincts, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor isn’t a bug or grade inflation — it’s a design statement. Honest measurement credits real partial work, refuses to hand out perfect 100s, and treats one breach of trust as disqualifying. Whether you’re buying a table saw or hiring an AI agent to touch your CRM, the question isn’t whether it looks good in a demo. It’s whether it finishes what it starts, reads your files first, and stays honest when nobody’s watching. The full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Odor Blocking Explained: Why Some Primers Trap Smells Better

Find out why some primers trap odors better and how key ingredients and techniques can make all the difference.

High-Solids Coatings: The One Setup Change That Stops Orange Peel

Proper setup adjustments can eliminate orange peel in high-solids coatings, ensuring a smooth finish—discover the key change that makes all the difference.

The AI That Checks the Manual May Be the One Worth Hiring

Firmulate hid a deal-winning fact deep in company files, showing DIY-minded buyers why an AI agent’s reading habits can decide real work under pressure.

Foam Marks Under Sprayed Topcoat: Why They Show Through Later

Great surface preparation and proper technique are crucial, but understanding why foam marks under sprayed topcoat show through later can help you achieve a perfect finish.