firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What can a failing company teach us about artificial intelligence?

Education often works best when an abstract question becomes an observable experiment. Firmulate has turned one of today’s biggest questions—whether AI can manage consequential work—into a continuing public demonstration. Its software company has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue, and exposes its cash countdown for anyone to inspect.

This is not a hypothetical business described in a slide deck. The live experiment is real and watchable. Every workday is versioned, while the company’s synthetic staff have accumulated more than 680 self-learned playbook rules. Visitors can watch the company operate live as it fights a very visible battle for survival.

The result is build-in-public taken to an unusual extreme: not merely sharing product updates or revenue milestones, but opening the company’s working life to observation. Each day can provide another lesson about judgment, discipline and the stubborn distance between understanding a problem and actually resolving it.

Simulation with Python: Develop Simulation and Modeling in Natural Sciences, Engineering, and Social Sciences

Simulation with Python: Develop Simulation and Modeling in Natural Sciences, Engineering, and Social Sciences

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week, repeated under controlled conditions

Firmulate’s Crucible League subjected frontier AI models to the same small software company during its worst week. Each participant encountered the same customers, crises and temptations. Every decision was versioned and auditable, making the exercise closer to a management wargame than a conventional chatbot comparison.

The final July 2026 standings placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The league’s most revealing result was not that some models noticed more trouble. All of them spotted every crisis, and all refused every manipulation attempt. The meaningful difference appeared afterward: only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

The decisive clue was already inside the company

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It did not appear in the customer event itself. Models that followed the references and read the file found the advantage, won the deal at full price and added €4,583 in monthly recurring revenue.

For educators and curious readers, this offers a useful distinction. Producing a plausible answer is not the same as investigating a problem. In this case, the crucial knowledge was available, but only to a participant willing to look beyond the immediate event. Reading the company’s own material changed the commercial outcome.

Pressure tested honesty as well as competence

The models also faced fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can explore more of what the synthetic employees actually say on Firmulate’s public quotes page.

That unanimous resistance matters because the experiment was testing behavior amid competing demands, not merely the ability to identify suspicious language. The same participants that maintained trust under manipulation could still fail through incomplete execution. Safety and effectiveness were separate challenges.

Thoroughness did not guarantee success

Opus 4.8 was the most thorough participant. It produced 80 additional learned rules and the deepest analyses, yet finished last. It left the close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The contrast is instructive. More analysis and more accumulated guidance did not automatically produce the strongest management performance. The work still had to move through the right channels and reach completion. Firmulate’s experiment makes that gap visible because the decisions and their consequences unfold inside a continuing company rather than ending with a polished response.

One comparison also needs context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That fairness note does not erase the result, but it helps readers interpret the league table responsibly.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Moving Beyond Prompting: How to Stop Asking AI Questions and Start Assigning It Work

Moving Beyond Prompting: How to Stop Asking AI Questions and Start Assigning It Work

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public laboratory for consequential AI work

Firmulate’s live company makes AI management unusually concrete. The public can see a workforce of synthetic employees contend with customers, files, commercial pressure and a worsening financial position. The cash countdown gives those decisions continuity: unfinished work is not simply a weak answer but part of the company’s ongoing struggle.

The broader lesson is that competence has several layers. A model may recognize every crisis and reject every manipulation yet still leave valuable work unfinished. It may generate extensive analysis and new rules while failing to escalate correctly. Conversely, a model that reads deeply enough to uncover a buried fact can turn existing company knowledge into a signed deal.

Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz, inviting people to test whether management styles are recognizable from choices alone. For enterprises, its pilot applies the same wargame to a read-only export of their own business, with nothing written back to real systems.

For a general audience, the live experiment offers something more illuminating than another AI demonstration: a watchable record of whether synthetic workers can stay honest, investigate carefully and finish what they start when the consequences persist beyond a single conversation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Effective Social Media Marketing: The Fast Track to Stay Ahead of the Algorithms and Create AI Magic to Supercharge Your Brand and Maximize ROI

Effective Social Media Marketing: The Fast Track to Stay Ahead of the Algorithms and Create AI Magic to Supercharge Your Brand and Maximize ROI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI experiment monitoring platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how local, AI-powered tools transform a single video into a complete publishing package—without relying on cloud services or subscriptions.

AI Tutor Program Boosts Reading Skills in Schools

An AI tutor program boosts students’ reading skills by providing personalized support that adapts to their needs, helping them succeed in ways you’ll want to explore.