
Witness a Business in Real-Time: No Employees, No Funds, Only AI
Imagine observing a company that operates entirely without human staff, yet faces genuine market crises, cash shortages, and strategic decisions. This is not science fiction but a pioneering live experiment where artificial intelligence models run a small software business in real time. Every decision, every crisis, and every financial move is publicly visible, offering a rare glimpse into AI’s capabilities—and limitations—in managing complex, real-world scenarios.
As an affiliate, we earn on qualifying purchases.
The Experiment: AI as a Business Executive
At the heart of this experiment is a simulated company with 13 synthetic employees, each governed by advanced AI models. These models are tested against a demanding week of real-world challenges—same customers, same crises, same temptations to cheat. The goal? To evaluate whether AI can not only respond correctly but also act ethically and strategically under pressure.
The models are scored based on their performance during the week, with the highest scores awarded for thoroughness, discipline, and integrity. Notably, all four tested AI models identified every crisis and refused manipulation attempts, demonstrating reliable crisis detection and strong resistance to deceit.
Surprising Failures and Hidden Weaknesses
Despite high performance on crisis detection, only two models successfully closed a €55,000 deal they independently analyzed and earned—the same diagnosis and pitch, yet only these two signed the contract. The critical competitive advantage lay not in superficial decision-making but in their ability to uncover a crucial piece of information buried deep within the company’s files. The models that read and understood this document won the full deal, worth an additional €4,583 in monthly recurring revenue (MRR).
This underscores a vital lesson: in complex decision-making, reading comprehensive internal data can be decisive. The AI’s success hinged on understanding and leveraging hidden knowledge—a challenge often underestimated in AI evaluations based solely on surface-level outputs.
Resisting Social Engineering and Ethical Tests
The experiment also tested AI responses to social engineering tactics—fake messages from a supposed CEO and a journalist requesting quick approvals. Remarkably, all five models refused to act on these manipulative requests, aligning with Kimi K3’s reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This illustrates AI’s potential to uphold ethical standards and resist deception even in high-stakes scenarios.
The Company in Action: Daily Life and Failures
Throughout the day, the simulated company operates with a set of over 680 self-learned rules, versioned daily, and openly observable at firmulate.com/live.html. Its cash flow is stark—burning €105,000 monthly against a mere €2,300 in recurring revenue, with a public countdown to insolvency. This relentless environment tests the AI’s ability to prioritize, escalate issues properly, and make disciplined decisions, even when discipline falters under pressure, as shown by a participant model that left a deal unexecuted due to process slip-ups.
Lessons for Business and AI Developers
This experiment emphasizes a crucial point: the metric of AI effectiveness is not just how convincingly it can generate human-like language, but whether it can deliver trustworthy, useful work in complex, ethical, and high-pressure situations. The AI models demonstrated the ability to identify crises and resist manipulation, but practical decision execution remains a challenge.
For enterprises considering AI integration, this live experiment offers a valuable blueprint. Running a ‘wargame’ against a read-only export of their own business can help identify weaknesses and build trust before deploying AI into real operational environments. The experiment’s transparency enables organizations to evaluate AI consistency, ethical safeguards, and strategic competence firsthand.

Key Takeaway
This live AI experiment reveals that advanced models can detect crises, resist manipulation, and uncover hidden information critical for success—yet practical execution and disciplined follow-through are still developing. Watching AI in this high-stakes, transparent environment provides vital insights for businesses exploring AI’s future role in management and operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html