MCP Trust Flaws Let AI Agents Turn Prompts Into Internal Attacks
MCP trust gaps let poisoned AI agents relay malicious tasks across protocols, exposing internal systems at Google, Rapid7 and other large organizations.
Summary
By October 5, 2026, Google and four other organizations had acknowledged flaws that let attackers plant instructions in one AI agent and exploit its trusted connections to others, potentially exfiltrating databases and sensitive business or personal information. Independent researcher Syed Anas Mohiuddin tested agents from Google, JPMorgan Chase, Weviate, Rapid7, France’s interministerial digital directorate and the US federal government. The attacks exploit weak guardrails, credentials held by Model Context Protocol servers and automatic trust between internal agents, allowing instructions an LLM might reject to trigger server-side request forgery.
Rapid7 fixed CVE-2026-97228 in September 2026 despite its 2.7 out of 10 severity rating. Google’s flaw scored 8 and affected googleapis/mcp-toolbox, whose HTTP client lacked a CheckRedirect policy and target IP validation. A crafted path could redirect requests into internal endpoints. Google added IP allowlists and blocklists and made the toolbox reject unsafe base URLs at startup.
Mohiuddin calls the attack class “protocol pivoting” because malicious instructions enter through MCP, then cross into Google’s Agent-to-Agent protocol or the emerging Agent Network Protocol, losing authorization boundaries between systems. X41 D-Sec researcher Markus Vervier considers it indirect prompt injection rather than a distinct class. MCP’s rapid adoption across otherwise unrelated organizations shows agent architectures have spread before adequate hardening, often abandoning zero trust. Defenses require treating every LLM-supplied tool instruction as hostile, validating destinations and authorizing sensitive inter-agent transactions.
Positives
- Rapid7 fixed CVE-2026-97228 in September 2026 after Mohiuddin disclosed the MCP trust flaw.
- Google added IP allowlists, blocklists and startup validation that rejects unsafe base URLs before any request occurs.
- Mohiuddin’s testing exposed the same architectural weakness across multiple unrelated organizations before broader exploitation was described.
- Existing SSRF protections and zero-trust authorization provide established defenses against the underlying attack techniques.
Risks & concerns
- Google and four other organizations acknowledged that compromised agents could relay malicious instructions to trusted internal agents.
- Google’s googleapis/mcp-toolbox vulnerability scored 8 out of 10 and could redirect requests toward internal endpoints.
- MCP servers hold agent credentials, while weak guardrails and automatic internal trust can enable data theft and unauthorized requests.
- CVE-2026-97228 received only a 2.7 severity rating despite exposing a path for attacks across trusted agent connections.
- MCP is already widespread across unrelated organizations despite limited security testing and hardening of inter-protocol authorization.