PofoliaPofolia ile paylaşıldı

ACM Transactions on Multimedia Computing Communications and Applications· 2024Q1

İşaret Dili Üretimi için Gloss Odaklı Koşullu Difüzyon Modelleri

Gloss-driven Conditional Diffusion Models for Sign Language Production

Shengeng Tang, Feng Xue, Jingjing Wu, Shuo Wang ve diğerleri

Kısa özet

Yeni bir Gloss Odaklı Koşullu Difüzyon Modeli (GCDM), işaret gloss dizilerini bir Transformer ile kodlayıp çapraz dikkat yoluyla bir difüzyon modelinin gürültüsüne entegre ederek metinden işaret dili videoları üretir.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Sign Language Production (SLP) aims to convert text or audio sentences into sign language videos corresponding to their semantics, which is challenging due to the diversity and complexity of sign languages, and cross-modal semantic mapping issues. In this work, we propose a Gloss-driven Conditional Diffusion Model (GCDM) for SLP. The core of the GCDM is a diffusion model architecture, in which the sign gloss sequence is encoded by a Transformer-based encoder and input into the diffusion model as a semantic prior condition. In the process of sign pose generation, the textual semantic priors carried in the encoded gloss features are integrated into the embedded Gaussian noise via cross-attention. Subsequently, the model converts the fused features into sign language pose sequences through T-round denoising steps. During the training process, the model uses the ground-truth labels of sign poses as the starting point, generates Gaussian noise through T rounds of noise, and then performs T rounds of denoising to approximate the real sign language gestures. The entire process is constrained by the MAE loss function to ensure that the generated sign language gestures are as close as possible to the real labels. In the inference phase, the model directly randomly samples a set of Gaussian noise, generates multiple sign language gesture sequence hypotheses under the guidance of the gloss sequence, and outputs a high-confidence sign language gesture video by averaging multiple hypotheses. Experimental results on the Phoenix2014T dataset show that the proposed GCDM method achieves competitiveness in both quantitative performance and qualitative visualization.

Yazarların özeti; kaynağından alınmıştır. ACM Transactions on Multimedia Computing Communications and Applications, 2024 · DOI ↗

ÇıkarımlarUygulamada
Ana noktalarUygulamada
Makaleye SorUygulamada

Devamı Pofolia uygulamasında

Çıkarımlar, ana noktalar ve makaleye soru sorma; ilgi alanına göre her gün yeni özetler. Ücretsiz.

Web'de giriş yaparak aç

Alan: İnsan-Bilgisayar Etkileşimi

Human-Computer InteractionComputer Science