Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

About

Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short clips and applied to long-horizon videos through chunked sliding-window inference, often causing temporal discontinuities at chunk boundaries. We address this by reframing long-horizon relighting as \emph{temporally conditioned latent domain translation}. Our framework enforces cross-chunk continuity by propagating target-domain latents across boundaries and makes this behavior learnable using \emph{masked target-domain self-conditioning}, training the model to continue from temporally masked propagated context. We further introduce \emph{warm-start prompting} with a relit prompt anchor from a controllable generative model, which establishes the initial target-domain state and creates a general interface for prompt-based relighting. Experiments on in-the-wild long-horizon videos show markedly improved temporal consistency, with chunk-boundary artifacts largely reduced and unwanted appearance changes across chunks greatly suppressed.

Jing Yang, Mayoore Jaiswal, Zian Wang, Steven Zeng, Rochelle Pereira, Yajie Zhao, Jianyuan Min• 2026

Related benchmarks

TaskDatasetResultRank
Target-light reconstructionMulti-Illumination (held-out)
PSNR14.39
7
Video RelightingSynthetic dataset
Normal Map Error0.0971
4
Showing 2 of 2 rows

Other info

Follow for update