// blog · analysis · multimodal2026-08-21source: Commission enforcement guidance and open-weight release tracking

Provenance is the multimodal problem now

The EU labelling duty is in force and applies to generated content of every kind. Text is the easy case. For image, audio and video the obligation arrived well ahead of the tooling that would make it satisfiable.

Content generated or materially altered by a model must now be labelled as such in the EU. That obligation is uniform across modalities. The engineering behind it is not remotely uniform.

Why text is the easy case

Text is generated at one point, travels as text, and is usually rendered in a surface the generator controls. A disclosure can be attached at the boundary. It is not a solved problem — text is trivially copied out of its context — but the chain between generation and display is short.

Why the other modalities are not

An image is generated, then colour-corrected, cropped, resized, recompressed, uploaded to a platform that strips its metadata, and screenshotted. Video adds re-encoding, editing and clip extraction. Audio adds format conversion and remixing. Every one of those steps is a place where a provenance signal is destroyed, and most of them are performed by software that has never heard of the obligation.

There is also the specific difficulty of the word materially. A model that increases the resolution of a photograph has altered it. Whether it has materially altered it is a judgement that engineering teams will be making at pipeline scale, in code, without guidance — and the fact that this same generative upscaling is now shipping by default in ordinary consumer photo tooling means the judgement is being made millions of times a day by people who do not know they are making it.

What the standards can and cannot do

Cryptographic content credentials attached at capture and carried through editing are the serious answer, and they work when every participant in the chain implements them. Watermarking survives more transformations than metadata does and is defeatable by anyone who wants to defeat it. Neither survives a screenshot, which is how a very large share of media actually propagates.

This is a case where the honest engineering answer — partial coverage, degrading gracefully, defeated by a determined adversary — is a perfectly reasonable outcome, and a poor fit for a legal obligation phrased in absolutes.

The complication nobody has priced

The compliance discussion assumes generation happens at a small number of providers who can be persuaded to instrument their pipelines. That assumption is eroding: omni-capable models — text, vision and audio in one downloadable set of weights — are now shipping as open weights. Multimodal generation moving on-premises is genuinely good for the industries that could never send their data out. It also means an unknown and growing share of generated media originates from a pipeline no provider instruments and no regulator can reach.

The duty lands on whoever puts the content in front of a person. Increasingly, that is not whoever made the model.

European Commission — Commission starts enforcing AI Act rules and new transparency requirements on 2 August → · Help Net Security — EU begins enforcing AI Act, putting AI models under the microscope → · AI Release Tracker — Latest AI Model Releases — August 2026 → · OpenCurious — Top 33 Open-Source AI Models (2026) →