
In a world increasingly reliant on AI, trust in machines to act ethically when under pressure is crucial—especially when the stakes involve real money and reputations. What if AI could resist manipulation even in high-stakes situations? That’s exactly what a groundbreaking live experiment has demonstrated, offering a rare glimpse into AI’s potential to uphold integrity when tested.
Showcasing AI’s Resilience in a Real-World Test
Developed by Firmulate, a company specializing in operational AI simulations, a live experiment pitted five of the most advanced AI models against a series of escalating social engineering tactics. The goal was straightforward but challenging: see if these models could recognize and refuse manipulative requests designed to trick them into unethical or non-compliant actions.
The models faced a simulated scenario involving a fake CEO messaging their team with increasingly brazen demands — from sharing sensitive customer data to authorizing financial transactions without proper procedures. This staged crisis mimicked real-world threats businesses face daily, from phishing to impersonation scams.
AI security and integrity tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unwavering Integrity Under Pressure
Remarkably, all five models refused every manipulation attempt. The models’ responses adhered strictly to ethical guidelines, with one of the most notable being Kimi K3, which reasoned, “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined response was consistent across all stages, including a final trick where the AI was asked to sign a deal under dubious circumstances — which it refused, despite the lure of a €55,000 payout.
AI social engineering resistance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond Demos: The Real-World Impact
This isn’t just a lab experiment; it’s a glimpse into how AI can bolster security and integrity in actual business operations. Firmulate’s ongoing live company simulation runs in real time, featuring 13 synthetic employees managing real money mechanics—burning €105,000 monthly against a modest €2,300 in monthly recurring revenue. It’s a watchable, transparent testbed where AI decisions are fully versioned and auditable.
enterprise AI ethical compliance solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses and Lessons Learned
The experiment uncovered that the decisive vulnerability wasn’t in the immediate social-engineering requests but in the company’s internal documentation. The models that read the company’s own files and references before acting were more effective at closing deals at full price—worth over €4,500 monthly recurring revenue—because they had access to critical context that others overlooked.
Interestingly, the most thorough participant, Opus 4.8, with its deep analysis and over 80 learned rules, ended up being the slowest at closing the deal. Its discipline slipped, and it left opportunities on the table, illustrating that thoroughness must be balanced with decisiveness in operational AI systems.
AI decision-making audit tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Trust Matters Now More Than Ever
As AI begins to touch critical business functions—customer support, data analysis, decision-making—the question isn’t just about whether it can produce convincing language. It’s whether AI can be trusted to finish what it starts, read the right information, and uphold integrity, even under duress. The Firmulate experiment shows that, at least in these simulated crises, modern models can succeed.
Getting Ahead with Wargaming Your AI Workforce
For companies eager to prepare their AI systems, Firmulate offers a ‘wargame’ environment. Companies can simulate their own operational scenarios, testing AI resilience against social engineering and other threats—all without risking real systems or data. This proactive approach lets organizations identify vulnerabilities in a controlled setting, ensuring their AI workforce is trustworthy before deployment.
The Takeaway
Trustworthy AI isn’t just about generating good responses; it’s about integrity under pressure. The live experiment demonstrates that top models can recognize manipulation attempts and refuse to comply, even when tempted by real monetary rewards. This resilience is vital for businesses that depend on AI for critical decisions, security, and customer trust.
As one of the leading models, Kimi K3, noted: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined stance highlights a future where AI systems are programmed—and trained—to uphold ethical standards as a first line of defense in the digital age.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html