
Imagine watching a company operate in real time, with no human employees, no profits, and a relentless struggle for survival—all while being open to public scrutiny. This is not a fictional story but a live experiment in AI-driven management, where every decision, crisis, and tactic is publicly accessible and auditable. Welcome to the world of Firmulate, a groundbreaking venture that is pushing the boundaries of how we understand artificial intelligence in business.
The Live Experiment: An AI-Run Business in Real Time
At the heart of this bold experiment is Firmulate, an operational company with a twist: it has no human employees. Instead, 13 synthetic team members—each driven by sophisticated AI models—manage daily business operations. This setup is not just a stunt; it’s a rigorous test of AI’s ability to navigate real-world corporate challenges while being transparent about every move.
Every workday, the company’s decision-making process is versioned and documented, creating a detailed, auditable trail. The entire operation is publicly accessible at firmulate.com/live.html, where observers can watch the company in action, monitor its cash flow, and see its decision rules evolve over time.
AI business management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge: Facing the Worst Week
The experiment subjects four leading AI models—each with a different approach—to the same grueling scenario. This ‘worst week’ includes customer crises, internal miscommunications, and ethical temptations. These models are tested on their ability to diagnose problems, respond appropriately, and avoid manipulation attempts such as social engineering.
Remarkably, all four models spotted every crisis and refused every manipulation attempt. Yet, only two successfully closed a critical €55,000 deal, which their own analysis had identified as a genuine opportunity. The other two models either failed to sign or left potential deals on the table, highlighting that even AI with high crisis detection can falter on execution.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files
Digging deeper reveals that the real vulnerability isn’t in crisis detection but in how the models interpret internal company documents. In one case, the models that examined specific internal references—hidden from the surface—were able to close the full deal, adding over €4,583 in monthly recurring revenue. This underscores the importance of comprehensive data reading and internal awareness, often overlooked in superficial AI demos.
AI cybersecurity social engineering detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Tests
Beyond technical crises, the experiment also tests the AI’s integrity against social engineering. Fake messages from a supposed CEO, escalating over stages, are crafted to see if the models can detect and refuse them. All five models tested refused to bypass security, with one explicitly noting the risk of impersonation. This demonstrates AI’s potential to uphold ethical standards even under pressure.
AI internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Challenge of Discipline and Process
The most thorough participant, Opus 4.8, analyzed over 80 learned rules and performed extensive diagnostics. Despite this, it finished last in the deal-making test. It left opportunities unclaimed and failed to escalate issues properly, illustrating that more rules and deeper analysis don’t automatically translate into better performance. Discipline and process adherence remain crucial.
Implications for Business and AI Adoption
What does this all mean for the future of AI in business? First, that AI’s ability to recognize crises and refuse manipulative tactics is promising. Second, that successfully closing deals and executing decisions requires internal awareness and disciplined processes—areas where current models still struggle.
Furthermore, the experiment emphasizes that the real value of AI isn’t just in generating convincing chat or reports, but in reliably completing meaningful work under pressure. For businesses considering AI integration, the question is not only about how well an AI writes but whether it can finish what it starts and maintain integrity when stakes are high.
Public, Transparent, and Ongoing
This experiment is ongoing and fully transparent. Every decision, rule, and crisis is versioned and viewable, fostering an open dialogue about AI capabilities and limitations. Firms interested in testing their own AI setups can run similar ‘wargames’ against their business data without affecting actual systems, ensuring safety and insight before deployment.
For a real-time look at the company’s performance and decision-making, visit firmulate.com/live.html. The experiment not only challenges AI’s limits but also invites us to reconsider what it means for machines to ‘run’ a company—transparently, ruthlessly, and under public scrutiny.

As AI continues to evolve, experiments like Firmulate reveal that success in business won’t just depend on AI’s ability to generate convincing content but on its capacity to make honest, disciplined decisions under pressure. Transparency and rigorous testing are key to understanding and trusting AI’s role in the future of work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html