A benchmark for attacking image-to-video models
VPA-Guard proposes defending and benchmarking image-to-video generation against visual prompt attacks — the first serious treatment of an input channel that has been shipping in products for a year.
VPA-Guard sets out to defend and benchmark image-to-video generation against visual prompt attacks: adversarial content placed in the input image rather than the text prompt.
The attack surface has been obvious and unaddressed for a while. Text prompts get filtered, logged and moderated because the industry has spent three years learning to distrust them. An input image arrives through a path built for content, not for instructions — and a model that reads instructions from pixels will read them from an image someone else supplied.
Publishing a benchmark alongside the defence is the right order of work. Defences without a shared measurement produce a literature of incomparable claims, each evaluated against whatever attack its authors imagined. A benchmark makes the next paper's numbers mean something.
The timing is behind the deployment, which is the usual pattern and still worth saying. Image-to-video has been in consumer products for a year; the security treatment of its input channel is arriving now.
It is the same shape as the agent story this month: capability standardises first, and the question of what a hostile input can make the system do gets its own paper afterwards.
arXiv — VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks → · arXiv — Seedance 2.0: Advancing Video Generation for World Complexity →