// news · multimodal2026-08-02source: digitalapplied / manilatimes

ByteDance's Seedance 2.0 unifies audio and video generation as Qwen previews a trillion-parameter multimodal model

ByteDance's Seedance 2.0 supports text, image, audio, and video inputs in a unified audio-video architecture, generating synchronised audio and video from 4 to 15 seconds at native 480p and 720p. At the World AI Conference in Shanghai, Alibaba's Qwen team previewed Qwen3.8-Max — its first multimodal model above a trillion total parameters, processing text, images, video, and documents.

Seedance 2.0's unified audio-video architecture is the technically ambitious part. Generating synchronised sound and picture from a single model — rather than producing video and dubbing audio separately — means the model has learned the joint distribution of what a scene looks and sounds like, which is a harder and more useful object than either modality alone.

Qwen3.8-Max crossing a trillion parameters as a multimodal model is the scale statement. Processing text, images, video, and documents in one model at that size signals that the Chinese labs are pursuing the same all-modalities-in-one-model frontier as the Western leaders, and doing it at the parameter counts that used to be the preserve of the largest language models.

Together they mark the multimodal frontier as a genuinely global, fast-moving contest. A Chinese short-video company shipping unified audio-video and a Chinese cloud lab previewing a trillion-parameter omni-model, in the same window as FLUX 3 and Muse in the West, is the multimodal field converging on the same target from every direction at once.

See our analysis →

Digital Applied — Seven days, seven model releases: the new AI normal → · The Manila Times — Vivideo.ai expands AI video platform with unified multi-model workflow →