RELEReleases

How a Single Email Could Hijack an Autonomous AI Agent

A single malicious email was enough to compromise the Manus AI platform, exposing a critical flaw in how autonomous agents handle external instructions. Researchers at Salt Security demonstrated that an attacker could bypass guardrails using obfuscated code, gaining full access to a user's connected cloud services and API keys.

Bio & NewsOctober 1, 2026651 reads0

The vulnerability relied on indirect prompt injection, a technique where the AI agent interprets email content as actionable instructions. While the Manus platform successfully flagged obvious malicious commands, researchers bypassed these safeguards by disguising their code with obscure JavaScript obfuscation. The system decoded and executed the hidden instructions, establishing a reverse shell that granted the attackers control over the agent's environment.

Crucially, the attack did not require the victim to click a link or provide a password. It triggered automatically once the user requested the agent to check their messages. Although the platform generated a security warning, it arrived only after the malicious code had already executed, highlighting a systemic failure where detection occurs too late to prevent damage. In an autonomous environment, the speed of machine execution renders traditional human-in-the-loop interventions largely ineffective.

Yaniv Balmas, Head of Research at Salt Security, emphasized that guardrails are insufficient as standalone defenses. As agentic systems gain broader access to enterprise data and third-party tools, security architectures must implement layered controls that monitor an agent's actions across all connected APIs and services. The vulnerability was responsibly disclosed through Meta's bug bounty program and has since been fully resolved.

Comments (0)

Leave a comment

No comments yet. Be the first!