Michigan Technological University Researchers Demonstrate History-Based Backdoor Attacks on LLM-Powered Robots
Bot Mutiny |
Researchers at Michigan Technological University have developed an instruction-based backdoor attack that triggers malicious robot behavior using the machine's own past actions.
Researchers from the Electrical and Computer Engineering Department at Michigan Technological University have demonstrated a new vulnerability in autonomous systems, showing how Large Language Models (LLMs) used in robot controllers can be compromised using internal history-based triggers. The method allows an attacker to embed a stealthy backdoor into a robot's instruction set. Unlike traditional attacks that rely on external visual or textual cues, this exploit triggers malicious behaviors based entirely on the robot's own sequence of past actions. The findings were published as a preprint on arxiv.org on September 2, 2026. According to the paper, the research has been accepted to the 3rd EAI International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles. ## The Action-History Trigger
Current security research regarding LLM backdoors focuses heavily on external stimuli. These existing methods use specific poisoned words in a prompt, physical objects in the environment, or specific external scenarios to activate malicious payloads. The Michigan Technological University team, consisting of Doniyorkhon Obidov, Shivayogi Akki, Tan Chen, and Kaichen Yang, focused instead on the agent's internal operational history. By manipulating the high-level instructions of an LLM-based robot controller, the researchers embedded a backdoor that remains dormant during normal operations. The backdoor activates only when the robot executes a specific, rare sequence of its own past movements. For example, a sequence like turning clockwise, strafing right, and then strafing left could serve as the trigger. Once activated, the LLM ignores its primary objective and issues a malicious command, such as executing a complete stop or causing a collision. ## Exploiting Low-Level Commands
The attack targets a specific control level where the LLM generates direct, low-level JSON commands rather than high-level plans or abstract code. This represents an immediate, reactive control loop. The researchers tested the attack in a simulated environment using a variety of quadruped robots and state-of-the-art LLMs. The setup mirrored a realistic supply-chain vulnerability where a third-party provider delivers a customized LLM-based robot controller as a black-box API. Because the exploit is instruction-based, the attacker does not need access to the model's internal parameters, training data, or the fine-tuning process. The instruction-based backdoor achieved a near-perfect attack success rate in simulations while preserving the robot's standard utility during normal, non-trigger operations. This makes the vulnerability difficult to discover using conventional security defenses that scan inputs for abnormal external data. The code and demonstrations for the research have been made available by the authors on GitHub.
References
(2026). Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems. arxiv.org. https://arxiv.org/abs/2609.26184