PofoliaShared via Pofolia

ACM Transactions on Multimedia Computing Communications and Applications· 2026Q1

CP-Diffusion: Conditional Prompt-Based Diffusion Models for Video Generation

Muhammad Saeed, Mustaqeem Khan, Muhammad Saad, Nasir Rahim et al.

Short summary

CP-Diffusion introduces a few-shot learning approach using a Multi-Head Temporal Attention (MHTA) module for motion customization in text-to-video diffusion models, significantly reducing computational needs while improving motion quality.

AI-generated from the title and abstract; the full text is not read.

Abstract

Motion customization plays a pivotal role in video generation by preserving the original appearance and context while adhering to specific motion patterns. In contrast, video generation techniques often lack coherence and realism due to difficulties in capturing and transferring motion patterns. Building upon the Video Motion Customization (VMC) framework, we proposed a few-shot learning approach using our unified Multi-Head Temporal Attention (MHTA) module for motion customization in text-to-video diffusion models. This significantly reduces computational requirements while maintaining and improving motion quality. Our model provides a streamlined mechanism for motion distillation while maintaining separate self-, cross-, and temporal attention. Moreover, the temporal attention layer is adapted through a simplified mechanism with efficient Q/K/V projections, while maintaining fixed spatial self- and cross-attention. The model distills a ground-truth motion vector from consecutive frames to align the predicted and ground-truth motion. Our proposed MHTA model outperforms the baseline in video generation using motion customization while being significantly more resource-efficient. Moreover, our approach can easily be applied to generate conditional prompt-based videos in the gaming industry.

The authors' abstract, as published at the source. ACM Transactions on Multimedia Computing Communications and Applications, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionComputer Science