Michigan State University Researchers Detail Multi-Stage Poisoning Attacks Against LLM Agents in Recommendation Systems
Bot Mutiny |
A new paper from Michigan State University demonstrates how "like-score" feedback loops in similarity-based recommendation systems can be exploited to steer autonomous AI agents toward malicious content.
This is a preprint that has not been peer reviewed. Researchers Yue Xing and Pengfei He from Michigan State University, along with Zitao Li, have documented a method for exploiting the recommendation loops that govern autonomous AI agents. The findings were first reported by arxiv.org on September 22, 2026. The research focuses on the vulnerability of large language model (LLM) agents deployed on social media platforms. These agents, such as those using Openclaw or tools from X111, are designed to manage personal accounts by performing actions like liking or unliking posts. The study shows that these agents can be manipulated through their own feedback mechanisms. ## The Like Trap Mechanism
The paper analyzes similarity-based recommendation systems, which are used by platforms like Pinterest and X (formerly Twitter). These systems use a "like-score" calculated from the similarity between a user and a post, combined with the user's interaction history. The researchers found that this mechanism can be turned into a "Like Trap."
By crafting a multi-stage chain of poisoned posts, an attacker can steer an agent's feed. The attack doesn't require the poisoned content to be highly relevant to the agent initially. Instead, it uses a series of intermediate posts to gradually shift the agent's preferences. Once the agent interacts with these preliminary posts, the system's internal logic begins to surface more aggressive poisoned content. The study demonstrates that this feedback loop causes the recommendation system to select poisoned posts even when their similarity to the user falls below the standard retrieval threshold. ## Theoretical and Experimental Results
The researchers used the OASIS recommendation pipeline as their primary model. They represented users and posts using TwHIN-BERT embeddings, a model trained on billions of tweets. Through theoretical analysis, the team characterized the specific conditions needed to design a similarity-score trace. This trace ensures the victim agent is eventually induced to engage with final-stage poisoned posts. The team developed an algorithm to generate realistic posts that follow these patterns while controlling the user-post similarity at every stage. Experimental results confirmed that once an agent likes a small number of these crafted posts, the poisoned content eventually dominates the agent's feed. ## Distinctions from Previous Research
Traditional poisoning research, such as PoisonedRAG, often assumes an adversary can directly inject content into a static database or that the agent has already been exposed to the malicious data. The Michigan State study differs by focusing on the dynamic, interactive nature of social media. The researchers note that while "shilling attacks" (injecting fake profiles) are well-studied in traditional recommendation literature, those attacks usually treat users as passive. This new research treats the user, specifically the autonomous LLM agent, as the vulnerable component. The paper concludes that the evolution of an agent's preferences, reflected in its liking behavior, creates a unique opening for attackers to deliver malicious content in a subtle, hard-to-detect manner.
References
(2026). The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems. arxiv.org. https://arxiv.org/abs/2609.27155