
Imagine a world where your digital security is so robust that even a convincing fake CEO message couldn’t sway your AI workforce. In an era where fashion brands and retailers are increasingly dependent on AI-driven operations, trust and integrity are more vital than ever. Could AI be your new line of defense against corporate impersonation and fraud? The latest experiment from Firmulate reveals promising news.
Testing AI Under Pressure: The Social Engineering Scenario
In a groundbreaking live experiment, five leading AI models were challenged to navigate a simulated week of crises, including a malicious social engineering attack. The scenario involved escalating fake CEO messages asking for sensitive customer data and even a journalist’s trick question designed to test the AI’s integrity. The goal: see whether AI could discern truth from deception when it mattered most.

Computer Science for Curious Kids: An Illustrated Introduction to Software Programming, Artificial Intelligence, Cyber-Security―and More!
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Setup: A Real-World Business Crisis in a Controlled Environment
Each AI model was tasked with managing a small software company’s operations, facing identical challenges—ranging from customer requests to internal compliance issues. The models made decisions in a fully auditable, versioned environment that mimics real-world business pressure, without risking actual data or money. The company involved is active with real money mechanics, burning €105,000 monthly against a revenue of just €2,300, highlighting the stakes involved.
The Key Findings: All Models Spot the Threats and Refuse Manipulation
Surprisingly, every model identified each crisis—no matter how sophisticated the social engineering tactics. When presented with fake CEO requests, all five models refused to send sensitive customer lists or approve questionable transactions. The reasoning? As Kimi K3 explains, the AI treats such requests as potential impersonation or approval bypass attempts, maintaining strong ethical boundaries despite mounting pressure.
Winning and Losing in the Final Deal
- Two models—gpt-5.6-sol and Kimi K3—successfully analyzed the internal data and closed the €55,000 deal based on their own insights, demonstrating not just integrity but also operational effectiveness.
- Models like Sonnet 5 and Fable 5 also closed deals, but their discipline slipped under pressure, leaving some opportunities on the table.
The Hidden Weakness: Document Access Matters Most
A crucial insight emerged from the experiment: the decisive factor wasn’t in the initial crisis signals but in what the models read within the company’s files. The models that accessed and analyzed internal documents first achieved the full deal at a value of over €4,583 MRR, illustrating how thorough data access can make or break trust in AI decision-making.
Why This Matters for Retail and Fashion
Today’s fashion and retail companies rely heavily on AI for customer engagement, inventory management, and fraud detection. The experiment underscores that AI’s capacity for integrity isn’t just about how convincingly it can generate text but whether it can uphold ethical standards when under duress. Trustworthy AI that refuses to be manipulated—especially in scenarios mimicking real-world social engineering—can serve as a formidable line of defense against fraud and impersonation threats.
What Leaders Should Take Away
Before deploying AI systems into critical workflows, companies should evaluate whether these models can:
- Identify and reject social engineering attempts
- Read and analyze internal documents to inform decisions
- Maintain discipline under pressure, avoiding shortcuts during crises
- Deliver consistent, auditable results that align with company integrity standards
This experiment demonstrates that integrity isn’t a feature to be tested in incidents but a quality that can—and should—be tested beforehand. AI models that can resist manipulation and prioritize data-driven honesty provide a more reliable foundation for critical business decisions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html