Skip to main navigation Skip to search Skip to main content

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

  • Kaifeng Gao
  • , Jiaxin Shi
  • , Hanwang Zhang
  • , Chunping Wang
  • , Jun Xiao
  • , Long Chen*
  • *Corresponding author for this work

Research output: Chapter in Book/Conference Proceeding/ReportConference Paper published in a bookpeer-review

Abstract

With the advance of diffusion models, today's video generation has achieved impressive qual-ity. To extend the generation length and facilitate real-world applications, a majority of video dif-fusion models (VDMs) generate videos in an au-toregressive manner, ie., generating subsequent clips conditioned on the last frame(s) of the previ-ous clip. However, existing autoregressive VDMS are highly inefficient and redundant: The model must re-compute all the conditional frames that are overlapped between adjacent clips. This issue is exacerbated when the conditional frames are extended autoregressively to provide the model with long-term context. In such cases, the compu-tational demands increase significantly (ie, with a quadratic complexity w.r.t. the autoregression step). In this paper, we propose Ca2-VDM, an efficient autoregressive VDM with Causal gen-eration and Cache sharing. For causal gener-atlon, it introduces unidirectional feature com-putation, which ensures that the cache of con-ditional frames can be precomputed in previous autoregression steps and reused in every subse-quent step, eliminating redundant computations. For cache sharing, it shares the cache across all denoising steps to avoid the huge cache stor-age cost. Extensive experiments demonstrated that our Ca2-VDM achieves state-of-the-art quan-titative and qualitative video generation results and significantly improves the generation speed. Code is available: https://github.com/Dawn-LX/CausalCache-VDM

Original languageEnglish
Title of host publicationProceedings of the 42nd International Conference on Machine Learning
Pages18550-18565
Number of pages16
Publication statusPublished - 2025
Event42nd International Conference on Machine Learning, ICML 2025 - Vancouver, Canada
Duration: 13 Jul 202519 Jul 2025

Publication series

NameProceedings of Machine Learning Research
PublisherML Research Press
Volume267
ISSN (Electronic)2640-3498

Conference

Conference42nd International Conference on Machine Learning, ICML 2025
Country/TerritoryCanada
CityVancouver
Period13/07/2519/07/25

Bibliographical note

Publisher Copyright:
© 2025 by the author(s).

Fingerprint

Dive into the research topics of 'Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing'. Together they form a unique fingerprint.

Cite this