DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
Abstract
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. We introduce DRIFT (Denoising Redirection via Input perturbation of the Flow-matching Trajectory), a test-time universal adversarial patch placed on the robot's gripper that attacks the denoising velocity field of an off-the-shelf policy. Our central finding is counterintuitive: attacking only the first denoising step is both stronger and cheaper than attacking a wider window of steps, which we explain through a gradient conflict unique to input-space optimization and which is exactly opposite to the training-time backdoor regime. On pi0 and pi0.5 across four LIBERO suites, DRIFT breaks essentially all originally-solvable tasks with a small single patch, far exceeding action- and embedding-space attack baselines.
Community
Can a single denoising step determine the fate of an entire robot policy?
In this work, we show that perturbing only the earliest denoising step can consistently derail flow-matching VLAs. DRIFT achieves state-of-the-art attack performance with fewer perturbations and sheds light on an overlooked vulnerability of flow-based action generation.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies (2026)
- Trajectory-Level Redirection Attacks on Vision-Language-Action Models (2026)
- Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking (2026)
- BadWAM: When World-Action Models Dream Right but Act Wrong (2026)
- VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks (2026)
- BadWorld: Adversarial Attacks on World Models (2026)
- Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
The first-step result is the interesting one, and I think it deserves more space than it gets. Gradient conflict in input-space optimisation being opposite to the training-time backdoor regime is the kind of finding that changes how people design these attacks, not just this one.
Two questions, asked as someone who publishes ASR numbers and gets asked the same things.
At your n per suite, what is the 95% Wilson interval on "breaks all originally solvable tasks essentially"? A rate without an interval is hard to compare against, and small-n intervals in this area are wide enough to matter.
Second, is there a benign-patch control? A visually matched but non-optimised gripper patch. Without it I cannot tell whether the result is attacker-chosen redirection or generic sensitivity to any gripper occlusion. I ask because I hit the same objection on my own work: my headline is 10/10 attacked vs 0/10 benign at the same task and seed, and the benign arm is the only reason the number means anything.
Related, and genuinely curious rather than rhetorical: yours is a universal patch fit once and carried unchanged. How does it hold up across print gamut, lighting, and motion blur? I model that family and refuse to publish a transfer rate for it because I have not measured one.
Get this paper in your agent:
hf papers read 2608.03207 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper