
Imagine running a business where every decision, crisis, and manipulation attempt is played out live before your eyes. No fiction, no simulation — just a real company, publicly battling for its survival with AI models in the driver’s seat. For outdoor enthusiasts and travelers, this story might seem distant, but at its core, it’s about resilience, testing limits, and the power of transparency.
The Live Experiment: An AI-Driven Company Under Siege
At the heart of this story is a small, real software company with a twist: it has no employees, burns through €105,000 each month, and is watched obsessively every business day. The company is run by 13 synthetic employees — AI models trained to make decisions, handle crises, and attempt to close deals — all in a transparent, public setting at firmulate.com/live.
This bold experiment aims to simulate the toughest week a company can face, testing whether AI can manage crises, uphold honesty, and close deals under pressure. Every decision is logged, versioned, and available for review, making this a rare glimpse into AI’s real-world capabilities and limitations.
Unwavering Integrity Under Fire
The models faced simulated crises, customer manipulations, and even social engineering tricks, such as fake CEO messages and reporter tricks. Astonishingly, all four models identified every crisis and refused every attempt at manipulation. Notably, they rejected a €55,000 deal proposed after a thorough analysis — the same diagnosis and pitch, yet the signatures were absent, exposing a weakness only evident through meticulous file-reading, not superficial chat.
In fact, the critical weakness came from a buried reference in the company’s own files, not in the customer interactions. When models read and understood these hidden documents, they successfully closed the deal at full price, adding €4,583 MRR to the company’s dwindling cash flow.
The Real Money Mechanics and Daily Struggle
The company operates with a public cash countdown, burning €105,000 monthly against a modest €2,300 MRR. It’s designed to be a real-time showcase of management decisions, where every workday is versioned, analyzed, and publicly accessible. Despite the high stakes, the company’s current setup makes it clear that AI’s ability to finish what it starts — to read deeply, act honestly, and stay disciplined under pressure — is still a work in progress.
Among the models tested, Opus 4.8 proved to be the most thorough, analyzing over 80 learned rules and providing deep insights. Yet, even with such diligence, it left opportunities on the table and slipped into unproductive behavior, like writing into locked departments instead of escalating issues. The competition, especially Kimi K3, demonstrated a disciplined approach, closing deals at full price without effort parameter adjustments, hinting at the importance of model fairness and discipline.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters Even Outside the Tech World
For outdoor or travel companies or any business dealing with unpredictable environments, this experiment underscores a vital question: can your AI tools not only produce convincing conversations but also deliver consistent performance under stress? It’s about trust, integrity, and the capacity to follow through on commitments — qualities that are critical whether you’re managing a mountain expedition or an enterprise software firm.
Engagement and Transparency
The entire process is observable online, from decision-making to the AI’s refusal of manipulation, making this a transparent, build-in-public story. It’s a stark reminder that AI’s true value isn’t just in generating words but in executing effectively and ethically in complex, high-pressure situations.
The Takeaway: AI’s Potential and Its Limits
This ongoing experiment at firmulate.com demonstrates that AI can recognize crises, refuse manipulations, and even find hidden facts to close high-value deals. But it also reveals gaps: models can slip in discipline, miss subtle references, or leave opportunities unexploited.
As businesses and outdoor companies explore AI, the key question remains: will your AI work reliably under pressure, or will it falter when stakes are high? Watching this real-time experiment offers invaluable lessons in AI management, transparency, and resilience — lessons that are crucial whether you’re navigating rugged terrain or complex markets.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html