Video DeltaNet speeds up livestream video generation with smarter AI attention

Photo of author

By Sophia Chen

Generating high-quality livestream videos using artificial intelligence is a complex and resource-heavy task. A newly published research paper introduces Video DeltaNet (VDN), a novel approach that makes this process significantly faster without sacrificing video quality. This advancement could help power smoother and more efficient AI-generated video experiences, from real-time streaming to virtual events.

Key Takeaways

  • Video DeltaNet combines two types of AI attention mechanisms—local Softmax attention and a new linear memory-based attention—to balance fine detail and long-range video context.
  • The method updates video frame information efficiently by processing spatial tokens together, reducing computational bottlenecks common in video generation.
  • VDN achieves a 14.5 times speedup in generating 14.3-second, 768p videos compared to previous models, using eight NVIDIA B200 GPUs.
  • The approach integrates smoothly with existing pretrained models via a staged training process, making it practical for real-world deployment.

At the heart of many AI video generators is a process called “attention,” which helps the model focus on different parts of the video to generate coherent visuals. Traditional “Softmax attention” captures detailed local interactions well but struggles with long video sequences because it requires heavy computation. On the other hand, “linear attention” is faster and scales better but often misses the subtle details needed for high-quality video.

The researchers behind Video DeltaNet created a hybrid solution. They combined the strengths of local Softmax attention, which handles detailed frame-by-frame interactions, with a new technique they call “Video Delta Attention” (VDA). VDA uses a linear memory system that updates once per video frame by jointly analyzing all spatial tokens—essentially the small pieces of the image at each moment. This dual approach allows the model to maintain detailed visual quality while efficiently managing the broader context across frames.

To make this hybrid attention approach work with existing video generation models, the team developed a “staged teacher-alignment” training process. This method gradually introduces the new attention components into pretrained models, ensuring stable learning and high performance. They tested Video DeltaNet on a model called MiniMax H3 and demonstrated that it could denoise (improve) 14.3-second videos at 768p resolution in just 6.7 seconds using eight NVIDIA B200 GPUs. This is a remarkable speedup compared to the previous baseline, which took over 97 seconds for the same task.

Beyond faster processing, the researchers also optimized the serving infrastructure with a system called SGLang to support efficient deployment. This means Video DeltaNet is not just a theoretical improvement but is designed with practical applications in mind.

Looking ahead, this research could influence how AI-generated videos are produced in real time, making livestreams, virtual reality experiences, and other video applications smoother and more accessible. While the current work focuses on video-to-video generation and retains traditional attention methods for text and audio inputs, future developments may further unify these modalities for even more versatile AI video creation tools.

Based on research published on arXiv by Haocheng Xi, Yiming Xie, Hexu Zhao et al..

Editor's note

This article focuses on the confirmed update first, then points readers to the competitive and policy context that shapes the beat.

Article briefing

Generating high-quality livestream videos using artificial intelligence is a complex and resource-heavy...

Story details

  • Author: Sophia Chen
  • Published: September 20, 2026
  • Category: AI

Key developments

  • Generating high-quality livestream videos using artificial intelligence is a complex and resource-heavy task.
  • A newly published research paper introduces Video DeltaNet (VDN), a novel approach that makes this process significantly faster without sacrificing video quality.
  • At the heart of many AI video generators is a process called "attention," which helps the model focus on different parts of the video to generate coherent visuals.

Why this matters

This advancement could help power smoother and more efficient AI-generated video experiences, from real-time streaming to virtual events.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI