IEEE Transactions on Circuits and Systems for Video Technology· 2025Q1
Sıfır-Çekim Video Sentezi ve Düzenlemesi için Uzamsal-Zamansal Enerji Güdümlü Difüzyon Modeli
Spatio-Temporal Energy-Guided Diffusion Model for Zero-Shot Video Synthesis and Editing
- 12atıf
- Q1SCImago
- 2025yıl
Kısa özet
EnergyViD, yeniden eğitime gerek kalmadan çeşitli koşullar (metin, poz, stil vb.) altında sıfır-çekim video sentezi ve düzenlemesi sağlayan yeni bir difüzyon modelidir; önceden eğitilmiş ağları kullanarak genel enerji fonksiyonları tanımlar.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Diffusion-based generative models have exhibited considerable success in conditional video synthesis and editing. Nevertheless, prevailing video diffusion models primarily rely on conditioning with specific input modalities, predominantly text, restricting their adaptability to alternative modalities without necessitating retraining of modality-specific components. In this work, we present EnergyViD, a universal spatio-temporal Energy-guided Video Diffusion model designed for zero-shot video synthesis and editing across diverse conditions. Specifically, we leverage off-the-shelf pre-trained networks to construct generic energy functions, guiding the generation process under specific conditions without the need for retraining. To precisely capture temporal dynamics related to motion conditions (e.g., pose sequences), we introduce a novel kernel Maximum Mean Discrepancy (MMD)-based energy function, which minimizes the global distribution discrepancy between the conditioning input and the generated video. Our extensive qualitative and quantitative experiments demonstrate that our algorithm consistently produces high-quality results across a wide range of motion and non-motion conditions, including text, face ID, style, poses, depths, sketches, canny edges, and segmentation maps, in the context of zero-shot video synthesis and editing. We will release source code upon acceptance of the paper.
Yazarların özeti; kaynağından alınmıştır. IEEE Transactions on Circuits and Systems for Video Technology, 2025 · DOI ↗
Devamı Pofolia uygulamasında
Çıkarımlar, ana noktalar ve makaleye soru sorma; ilgi alanına göre her gün yeni özetler. Ücretsiz.
Web'de giriş yaparak açAlan: Bilgisayarlı Görü ve Örüntü Tanıma
Computer Vision and Pattern RecognitionComputer Science