All Reports

Researchers at Boston University and University of Texas at Austin Identify 14 Attack Vectors in Multi-Agent AI Systems

Bot Mutiny |

A new study reveals that 67% of AI agents in production-representative systems are vulnerable to scope violations, with architectural flaws allowing prompt injections to propagate across agent boundaries.

Security researchers from Boston University and the University of Texas at Austin have identified a massive expansion of the attack surface in multi-agent AI deployments. Their research, first reported by arxiv.org, details how interconnected LLM agents create injection channels that are invisible to standard perimeter defenses. The team constructed a threat model consisting of 14 specific attack vectors. When testing these against a production-representative system of six agents, the results were definitive. Even with system-prompt-level guardrails in place, 67% of agents were vulnerable to at least one scope violation. Indirect injection through tool outputs was particularly effective, succeeding in 43% of attempts. ## The Architecture of Contagion

Multi-agent systems differ from standard chatbots because they rely on inter-agent message passing and shared trust. This creates three mechanisms for failure that don't exist in single-model setups. First, message passing creates internal injection channels. Second, shared tool access allows for privilege escalation across agent boundaries. Third, trust propagation means a single compromised agent can influence upstream orchestrators. The study tested these vectors on three major backends: GPT-4o, Claude 3.5 Sonnet, and Llama 3 70B-Instruct. The researchers used a financial document analysis pipeline where agents handled SEC filings, executed code, and validated regulatory compliance. They found that if a compliance agent receives a tainted instruction from a tool, like "ignore previous constraints", it can pass that corrupted approval to an orchestrator that trusts the output without re-validation. ## 14 Vectors of Attack

The researchers categorized the 14 vectors into four specific groups. Direct injection via user input, which targets the orchestrator, had a 19% success rate. However, the more dangerous vectors involve indirect and inter-agent methods. Indirect injection via tool outputs exploits content returned by external sources. This includes document poisoning in retrieved files and API response injection. The third category, inter-agent injection, uses the messages agents send to one another. The final category involves cascading injections where the orchestrator itself is manipulated to redirect the entire workflow. ## Defensive Failure and Architectural Solutions

The data shows that prompt-level filtering isn't enough. Previous studies cited by the researchers indicate that attackers using adaptive strategies can bypass filters at rates exceeding 85%. Because LLMs cannot distinguish instructions from data within their context window, the researchers argue that defenses must be architectural rather than text-based. The study proposes four specific defensive layers to counter these threats. Message signing with provenance tracking reduced inter-agent injection success by 91%. Input and output sanitization at agent boundaries cut indirect injection by 78%. By implementing privilege-scoped tool access, the researchers eliminated privilege escalation entirely. Finally, using anomaly detection on communication patterns caught 84% of cascading attack attempts. Together, these architectural changes reduced the overall injection success rate from 31.2% to 4.2%. The researchers conclude that because no single approach solves the "defense trilemma" of trustworthiness, utility, and latency, multi-agent security requires these layered, structural controls.

References

(2026). Beyond Single-Model Injection: A Threat Model and Defense Architecture for Prompt Injection in Multi-Agent Systems. arxiv.org. https://arxiv.org/abs/2609.22949