DEV Community
Follow
Understanding Prism: Dynamic Sparse Attention for Joint Video-Audio Generation
Prism is a novel sparse-attention framework designed to make training high-resolution video and audio generation models more efficient. High-resolution video contains an immense number of visual tokens, making traditional dense attention computationally expensive due to its quadratic scaling. Audio adds another layer of complexity, requiring models to connect sounds to their visual origins. Prism tackles this by intelligently adapting its attention patterns based on the local content of the video.The framework divides video clips into spatiotemporal macro-zones, analyzing visual variations and audio-video cross-attention signals within each zone. Based on these signals, Prism dynamically shapes its attention blocks, focusing computational resources on areas with significant visual change or strong audio-visual correlation. This sparse attention approach avoids unnecessary computations on less informative parts of the video.A preview of Prism is available on Hugging Face, offering inference scripts for image-to-video and text-to-video-and-audio generation, as well as support for native joint video-audio training. The repository currently requires substantial GPU resources, with 720p inference demanding an 80GB GPU and higher resolutions necessitating multiple such GPUs. Native training for 1080p and 2K resolutions requires at least 32 or 64 80GB GPUs respectively.The authors report that Prism achieves up to 2.5 times faster training compared to full attention, along with improved generation quality, though these results are specific to their experimental setup. This research preview is intended for users comfortable with downloading checkpoints and configuring GPU environments. Prism's core innovation lies in its content-aware, dynamic sparse attention, which is particularly beneficial for training high-resolution joint video-audio models.