Researchers have developed a new method that makes AI-generated videos faster and more reliable without sacrificing quality. Video diffusion models, a cutting-edge type of artificial intelligence used to create videos from scratch, usually require many steps to produce clear, realistic results. This process can be slow and computationally expensive. The new approach, called Projected Distribution Matching Distillation (PDMD), speeds up video generation while reducing visual glitches and unnatural textures that often appear when trying to speed up these models.
Key Takeaways
- PDMD improves the training stability of video diffusion models, preventing common issues like oversaturation and visual artifacts.
- This method reduces the number of denoising steps needed to generate videos, making the process faster without extra computational cost.
- PDMD achieves higher quality scores on benchmark tests for both video and audio generation compared to previous distillation methods.
- The technique is simple to implement, requiring only a minor code change without additional training stages or data.
Video diffusion models create videos by gradually refining noisy images through multiple steps, called denoising evaluations. Typically, these models need tens of such steps to produce smooth and detailed videos, which can be slow and resource-intensive. To speed things up, researchers use a process called Distribution Matching Distillation (DMD), which trains a smaller or faster model (the “student”) to mimic a larger, slower one (the “teacher”) but with fewer steps. However, DMD can introduce errors during training, causing the output videos to degrade over time, showing unwanted artifacts and color saturation.
To tackle this, the authors of the new study analyzed why DMD training sometimes becomes unstable. They found that errors coming from a component called the “critic” — which evaluates how well the student model is learning — accumulate and cause problems in the student model’s updates. Their solution, PDMD, works by filtering out the part of the update that aligns with these critic errors. In other words, PDMD projects away the error component that leads to instability, keeping the model focused on the correct learning signals.
This projection technique is mathematically proven to remove a significant portion of the critic’s errors while preserving the useful information needed for learning. Importantly, this approach does not require extra computational overhead, additional training data, or complex changes to the model architecture. It is effectively a “one-line code fix” that can be added to existing DMD setups.
When tested on popular benchmarks, PDMD showed clear improvements. For example, on the Wan2.1 benchmark—a standard test for video generation quality—PDMD scored 83.73 with just four denoising steps, outperforming the original DMD by over one point. It also excelled in joint video and audio generation tasks, achieving the highest scores in both visual and audio quality compared to similar models with the same speed constraints. User studies further confirmed that videos generated with PDMD looked better, with more natural motion and clearer sound.
The implications of this work are promising for applications that rely on fast and high-quality video generation. These include video editing, special effects, virtual reality content, and even real-time video synthesis for entertainment or communication. By making video diffusion models more efficient and stable, PDMD could help bring advanced AI video tools to devices with limited computing power or accelerate workflows in creative industries.
Looking ahead, the researchers note that PDMD’s simple integration makes it easy to adopt in future video generation systems. Further exploration may extend this approach to other types of AI models that use distillation techniques, potentially improving the speed and quality of AI-generated content across different media formats.
Based on research published on arXiv by Zimo Wang, Junkun Yuan, Angtian Wang et al..
