// news · multimodal2026-08-05source: radicaldatascience / llm-stats

Multimodal safety classification now fits on a single consumer GPU

Shieldstral's 3 billion parameters cover both text and multimodal safety on one 16GB card while matching guard models up to seven times larger. The compute floor for responsible deployment just dropped to something a small team can own outright.

Cost has been the quiet reason safety tooling concentrated in a few hands. Running a large guard model on every request is expensive enough that most teams outsourced it, which meant a third party saw all traffic and set the boundaries. A classifier that fits on one card removes that constraint entirely.

Covering images and video in the same model matters more than it did a year ago. Synthetic media obligations now apply across modalities in the EU, and a text-only filter leaves the largest surface unguarded — which is precisely where labelling and watermarking rules are aimed.

The efficiency result deserves independent scrutiny before it is treated as settled. If a 3B model genuinely matches 21B guard models on both modalities, a substantial amount of moderation spend across the industry has been buying size rather than accuracy, and the correction will be quick.

See our analysis →

Radical Data Science — AI news briefs bulletin board for August 2026 → · LLM Stats — AI updates today (August 2026) →