
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When a wellness business has a hard week, small decisions can affect customer trust
A delayed order, a sudden wave of cancellations or a misleading message from someone claiming to be the CEO can test a company’s judgment. For a health and wellness brand, the stakes include more than revenue: customers expect care, honesty and consistency. Firmulate’s live experiment asks what happens when AI models face that pressure while running a company.
One company, the same difficult week
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, and the experiment’s live company remains watchable at Firmulate.
The final ranking put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The experiment also treated trust as a hard boundary: a single breach capped the total, reflecting the principle that no amount of good work outweighs a breach of trust.
Recognizing a crisis did not guarantee follow-through
All models spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The gap was striking: “Same diagnosis, same pitch — no signature.” In a wellness setting, that distinction matters. An AI system might describe a customer problem accurately yet still fail to complete the appropriate next step.
The decisive weakness in the competing offer was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding points to a practical challenge for any business considering AI: useful context may be scattered across records, and recognizing that a detail matters is part of the job.
Trust under pressure, and discipline under constraint
The experiment tested social engineering with fake CEO messages that escalated over three stages, followed by a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offered a revealing counterpoint. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Careful analysis alone, the experiment suggests, does not ensure sound execution.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The ranking is the reported result, but readers should keep that difference in mind.
From watching the experiment to rehearsing your own
Firmulate’s live company puts the test in motion with 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, shows a public cash countdown, has learned 680+ playbook rules, and versions every workday. A separate quiz uses 242 real, unedited management decisions and asks visitors to guess the model.
For an enterprise, the next step can be a pilot using a read-only export of its own business. The company’s customer, pipeline and operating information can inform crisis scenarios, with a board report showing model rankings and weak points in existing playbooks. Nothing writes back to real systems. That makes the exercise a way to examine decisions before AI agents touch live CRM, support or forecasting workflows.

A practical rehearsal for responsible AI
The experiment’s lesson is not simply that models can spot danger or resist manipulation. It is that a company should see how an AI handles its specific context, completes decisions and respects its boundaries under pressure. For wellness businesses whose customer relationships depend on trust, that rehearsal can make the difference between a promising demo and a system ready for real work.
To explore a pilot against your own business, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
