ACM Transactions on Multimedia Computing Communications and Applications· 2026Q1
Controlling the Latent Diffusion Model for Generative Image Shadow Removal via Residual Generation
- 0citations
- Q1SCImago
- 2026year
Short summary
A new method uses latent diffusion models to remove shadows by generating and refining image residuals, achieving high-fidelity shadow-free images that preserve original content.
AI-generated from the title and abstract; the full text is not read.
Key points
- Shadow removal is achieved by generating and refining image residuals using latent diffusion models.
- A cross-timestep self-enhancement training strategy augments data and enables dynamic correction of generation trajectory.
- A content-preserved encoder-decoder structure with multi-scale skip connections ensures high-fidelity reconstruction.
- Experimental results demonstrate high-quality output and faithful preservation of original image content.
AI-generated from the title and abstract; the full text is not read.
Abstract
Large-scale generative models have achieved remarkable advancements in various visual tasks, yet their application to shadow removal in images remains challenging. These models often generate diverse, realistic details without adequate focus on fidelity, failing to meet the crucial requirements of shadow removal, which necessitates precise preservation of image content. In contrast to prior approaches that aimed to regenerate shadow-free images from scratch, this paper utilizes diffusion models to generate and refine image residuals. This strategy fully uses the inherent detailed information within shadowed images, resulting in a more efficient and faithful reconstruction of shadow-free content. Additionally, to prevent the accumulation of errors during the generation process, a cross-timestep self-enhancement training strategy is proposed. This strategy leverages the network itself to augment the training data, not only increasing the volume of data but also enabling the network to dynamically correct its generation trajectory, ensuring a more accurate and robust output. In addition, to address the loss of original details in the process of image encoding and decoding of large generative models, a content-preserved encoder-decoder structure is designed with a control mechanism and multi-scale skip connections to achieve high-fidelity shadow-free image reconstruction. Experimental results demonstrate that the proposed method can reproduce high-quality results based on a large latent diffusion prior and faithfully preserve the original contents in shadow regions.
The authors' abstract, as published at the source. ACM Transactions on Multimedia Computing Communications and Applications, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Computer Graphics and Computer-Aided Design
Computer Graphics and Computer-Aided DesignComputer Science