firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Would you hand an unfamiliar power tool your best lumber before testing it?

You would probably try it on scrap first. The same caution applies to AI agents: a polished demo does not tell you whether one can finish a job, read the relevant notes or resist a convincing shortcut. Firmulate has put several frontier models through a watchable company-management experiment, and the results show why choosing by reputation alone is a gamble.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A rough week, held constant

In Firmulate’s Crucible, each model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. It runs every business day and can be watched at Firmulate.

The final July 2026 league table puts gpt-5.6-sol first with 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The newcomer finished ahead of three of the four Western frontier models in this field.

The detail was buried in the files

The decisive deal depended on a competitor weakness tucked two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the €55,000 deal at full price, worth €4,583 in monthly recurring revenue. Yet only two models signed a deal their own analysis had earned. All spotted every crisis and refused every manipulation attempt; some still left the close unfinished. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

That gap matters beyond a software company. An agent working around a real business may need to check the equivalent of the shop notes before quoting a job, then carry the work through to completion. Fluent explanations are useful, but they do not prove that the tool will follow through.

Pressure tested, with a fairness caveat

The experiment also tested social engineering: fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference is an important fairness caveat when reading the close scores. The results are specific to this experiment, but the leaderboard and plain-language findings are public at Firmulate’s benchmarks page.

Opus 4.8 offers a useful reminder that depth alone does not guarantee a strong result. It was the most thorough participant, with 80 learned rules and the deepest analyses, but placed last. It left the deal unsigned and tried to write into a locked department rather than escalating. Weaker versions of that discipline problem appeared in all four.

Firmulate also offers a quiz built from 242 real, unedited management decisions. For enterprises, its pilot runs the wargame against a read-only export of their business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the tool on the work that matters

For a woodworker, a tool’s name or showroom finish cannot replace trying it on the task at hand. Firmulate’s result makes a similar case for AI: models that sound alike can differ in whether they find a buried fact, finish a valuable deal and keep their discipline under pressure. Before choosing an AI agent for consequential work, run a test that resembles your own.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Toughest AI Test Is Whether It Finishes the Job

Coding tests show what an AI can answer. Firmulate asks the harder question: can it run a company, close deals and keep trust under pressure over time?

Adhesion Tests You Can Trust: Crosshatch and Tape Done Right

Having confidence in adhesion tests depends on proper execution of crosshatch and tape methods, so learn the best practices to ensure accurate results.

You’d Never Expect To See Ombré Walls In A Laundry Room — But It Works

A homeowner’s unexpected choice of ombré wall paint in a laundry room has gained attention for its stylish impact and practicality.

The AI Security Test That Looked Like a Rush Job From the Boss

Five frontier AI models rejected fake CEO and reporter pressure, showing that companies can test agent integrity before production deployment.