// news · multimodal · safety2026-08-19source: arXiv

A benchmark for attacking image-to-video models

VPA-Guard proposes defending and benchmarking image-to-video generation against visual prompt attacks — the first serious treatment of an input channel that has been shipping in products for a year.

VPA-Guard sets out to defend and benchmark image-to-video generation against visual prompt attacks: adversarial content placed in the input image rather than the text prompt.

The attack surface has been obvious and unaddressed for a while. Text prompts get filtered, logged and moderated because the industry has spent three years learning to distrust them. An input image arrives through a path built for content, not for instructions — and a model that reads instructions from pixels will read them from an image someone else supplied.

Publishing a benchmark alongside the defence is the right order of work. Defences without a shared measurement produce a literature of incomparable claims, each evaluated against whatever attack its authors imagined. A benchmark makes the next paper's numbers mean something.

The timing is behind the deployment, which is the usual pattern and still worth saying. Image-to-video has been in consumer products for a year; the security treatment of its input channel is arriving now.

It is the same shape as the agent story this month: capability standardises first, and the question of what a hostile input can make the system do gets its own paper afterwards.

See our analysis →

arXiv — VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks → · arXiv — Seedance 2.0: Advancing Video Generation for World Complexity →