
If you trust someone to care for your dog, you want more than a confident promise. You want to know how they respond when plans go wrong, a stranger asks for a favor, or the right choice takes more effort. Businesses face a similar test as AI takes on work: can it handle a difficult week, protect trust, and follow through?
Firmulate makes that question watchable. Its live experiment puts AI models in charge of a small software company and lets readers follow the decisions. The next step is more personal for a business: try a wargame using a read-only export of its own data.
A rough week, played out in public
In the final Crucible League, published in July 2026, five entries placed from first to fifth: gpt-5.6-sol with 95, Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The do-nothing baseline scored 26. The league’s trust standard is blunt: “no amount of good work outweighs a breach of trust.”
The experiment gave frontier models the same small software company, the same customers, crises, and temptations. Every decision was versioned and auditable. The test was not whether a model could produce a persuasive answer. It was whether it could steer the company through its worst week.

ukeetap Extra dicke Silikonmatte, wasserdicht mit erhöhtem Rand, 48,3 x 30,5 cm, BPA-frei, rutschfeste Hunde- und Katzenfuttermatte, Futtermatte für Futter- und Wassernäpfe, auslaufsichere Matte zum
- Premium Material: BPA-free, food-grade silicone
- 100% Leak-proof: High raised edges contain spills
- Easy to Clean: Smooth surface for quick cleaning
As an affiliate, we earn on qualifying purchases.
Seeing trouble is not the same as finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding, in Firmulate’s words: “Same diagnosis, same pitch — no signature.”
One clue sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The decisive gap was follow-through: a model could identify the right opportunity and even make the case for it, but still leave the close on the table.
The social-engineering test also put trust under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

MateeyLife Napfunterlage Katzen, Napfunterlage Hund Groß Hundenapf
- Safe & Less Mess: High-quality food-grade silicone material
- Raised Edge Design: 0.5-inch raised border prevents spills
- Non-slip & Waterproof: Keeps bowls in place and prevents mess
As an affiliate, we earn on qualifying purchases.
More thorough did not mean more effective
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four.
There is a fairness detail for readers weighing the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Hubulk Futtermatte für Hunde und Katzen, Größe L (48.3 x 30.5 cm), XL (61 x 40.6 cm) oder XXL (81 x 61 cm), 1.3 cm und 2.5 cm, erhöhter Rand, Silikon, rutschfest, groß (48.3 x 30.5 x 1.3 cm), grau
- Available Sizes: L, XL, XXL options
- Material: 100% Premium Food-Grade Silicone
- Non-Toxic: Hypoallergenic and BPA-free
As an affiliate, we earn on qualifying purchases.
From watching to trying it at your company
The live company is built from 13 synthetic employees, but its money mechanics are real. It burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its playbook has grown to 680+ self-learned rules, and every workday is versioned. Readers can watch the experiment at firmulate.com.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz at firmulate.com/quiz.html. The choices make the experiment more than a leaderboard: they let readers compare how the models handled actual decisions.
For an enterprise, the practical proposition is a pilot against its own business. Firmulate uses a read-only export to create a digital twin, then runs crisis scenarios and produces a board report with a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems. That offers a way to move from observing an AI company under pressure to examining how models might respond to your customers, pipeline, and rules.

Firmulate’s results suggest that recognizing a crisis and refusing a dishonest request are only part of the job. Models also need to find the evidence, make a sound decision, and carry it through without breaking trust. A pilot lets a business explore those behaviors against its own data in a read-only setting. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Gorilla Grip SilikonHaustierFuttermatte, wasserdicht, erhöhte Kanten, um Verschütten zu verhindern, einfache Reinigung in der Spülmaschine, für Hunde und Katzen, Platzierungstablett für Futter und
- Waterproof Design: 100% waterproof to protect floors
- Textured Surface: Keeps bowls in place and reduces spills
- Perfect Size: Measures 47 x 29.2 cm
As an affiliate, we earn on qualifying purchases.