
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
What if your AI could run a company — and actually make smart, honest decisions?
Imagine an AI managing your business during its toughest week, facing real crises, customer pressures, and ethical temptations — and doing it without a single slip-up. That’s exactly what a live experiment by Firmulate is testing, pitting four advanced AI models against each other in a real-world company scenario. The results might surprise you, revealing not just how well AIs spot problems, but whether they can be trusted to finish what they start.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The High-Stakes AI Business Wargame
At the heart of this experiment is a small software company experiencing its worst week — a perfect storm of customer issues, crises, and potential manipulations. Four frontier AI models, each with their own management style, oversee the company’s operations in real-time, making decisions that impact hundreds of thousands of euros. Every move is carefully versioned and auditable, creating a transparent window into how these AIs behave under pressure.
Measurable Management Personalities
Each AI model has a distinct personality: some are meticulous and thorough, others terse and to the point, and a few refuse to communicate noise or unverified requests. Scores from the latest Crucible League, a ranking of AI capabilities as of July 2026, show that GPT-5.6-sol leads with a 95, followed by Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77. The baseline score for a do-nothing approach is 26, underscoring how much progress these models have made in managing complex decisions.
Can They Spot the Critical Clues?
All four models identified every crisis and refused manipulation attempts, such as fake CEO messages or staged reporter tricks. Impressively, they maintained integrity even when escalation messages were sent in staged stages, with each model refusing to approve suspicious requests. Kimi K3, for example, explicitly treated such requests as potential impersonation, demonstrating a cautious, security-aware stance.
Who Wins the Deal?
The ultimate test was whether the AI could close a €55,000 deal by diagnosing the company’s issues correctly, presenting a compelling pitch, and following through. Only two models signed the deal — the ones that thoroughly read and interpreted the company’s internal files. The models that failed to dig deep or slipped into shortcuts did not close the deal, even though they had identified the same problems and presented similar solutions.
What Makes the Difference?
The key difference was reading depth. The model that found the buried facts within the company’s files, which were two document references deep, won the full-price deal, adding €4,583 monthly recurring revenue. This buried insight was invisible in superficial chat conversations but crucial for closing the business.
Does Management Style Impact Results?
Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and providing deep insights. However, it left the closing on the table and slipped into writing attempts that were stored in a locked department instead of escalating. Despite its depth, it performed the poorest in the final outcome. Meanwhile, Kimi K3 ran without an effort parameter, maintaining discipline and scoring just behind GPT-5.6.
Real Business, Real Money, Real Time
The experiment runs in a live, functioning company with 13 synthetic employees managing real cash flows, burning €105,000 monthly against a revenue of €2,300. Every decision is public, and the entire setup is accessible at firmulate.com/live. This transparency showcases how AI decision-making holds up under actual business conditions, not just idealized demos.
Why Should Business Leaders Care?
It’s not about chat quality or conversational flair. It’s about whether AI can truly finish what it starts: reading critical documents, staying honest under pressure, and making decisions that maximize value. As AI begins to touch CRM, support, or forecasting tools, the real question isn’t “Can it write well?” but “Will it stay disciplined and deliver results?”

The Bottom Line
In a live business simulation, AI models showed they can detect crises and refuse manipulation, but their ability to close deals hinges on reading deeply and maintaining discipline. Quality AI management isn’t just about clever responses — it’s about executing with integrity and thoroughness, especially when stakes are high.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.