PofoliaShared via Pofolia

IEEE Transactions on Circuits and Systems for Video Technology· 2025Q1

Spatio-Temporal Energy-Guided Diffusion Model for Zero-Shot Video Synthesis and Editing

L. Yang, Yikai Zhao, Zhaochen Yu, Bohan Zeng et al.

Short summary

EnergyViD is a new diffusion model that enables zero-shot video synthesis and editing across diverse conditions (text, pose, style, etc.) by using pre-trained networks to create generic energy functions, avoiding retraining.

AI-generated from the title and abstract; the full text is not read.

Abstract

Diffusion-based generative models have exhibited considerable success in conditional video synthesis and editing. Nevertheless, prevailing video diffusion models primarily rely on conditioning with specific input modalities, predominantly text, restricting their adaptability to alternative modalities without necessitating retraining of modality-specific components. In this work, we present EnergyViD, a universal spatio-temporal Energy-guided Video Diffusion model designed for zero-shot video synthesis and editing across diverse conditions. Specifically, we leverage off-the-shelf pre-trained networks to construct generic energy functions, guiding the generation process under specific conditions without the need for retraining. To precisely capture temporal dynamics related to motion conditions (e.g., pose sequences), we introduce a novel kernel Maximum Mean Discrepancy (MMD)-based energy function, which minimizes the global distribution discrepancy between the conditioning input and the generated video. Our extensive qualitative and quantitative experiments demonstrate that our algorithm consistently produces high-quality results across a wide range of motion and non-motion conditions, including text, face ID, style, poses, depths, sketches, canny edges, and segmentation maps, in the context of zero-shot video synthesis and editing. We will release source code upon acceptance of the paper.

The authors' abstract, as published at the source. IEEE Transactions on Circuits and Systems for Video Technology, 2025 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionComputer Science