Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SketchKeyAnime: Reference-anchored Sparse Key-Sketch Animation Synthesis

About

Traditional animation production relies heavily on manual drawing and iterative refinement, particularly for key-pose design, in-betweening, and character coloring. While existing animation and video generation methods have made notable progress, they typically depend on RGB boundary frames, dense frame-wise conditions, or complete sketch sequences, limiting their applicability under low-cost input conditions. We present SketchKeyAnime, a video diffusion framework for generating structurally controllable, appearance-consistent, and temporally coherent animations from sparse key-sketch inputs. Given a single reference RGB image and a few temporally indexed key sketches, SketchKeyAnime introduces a dual-branch conditioning mechanism to encode local geometric constraints alongside semantic-temporal context. It leverages Sketch Cross Attention to fuse reference image and sketch conditions with learnable gating, and incorporates an Adaptive Weighted Loss to strengthen supervision on key-sketch frames and line-art regions. Experimental results on the Aesthetic subset of Sakuga-42M show that our approach consistently outperforms representative animation interpolation and sketch-guided generation baselines. Compared to the best-performing baseline, SketchKeyAnime reduces EDMD by 31.9\% and FVD by 9.5\%, demonstrating superior sketch fidelity and temporal coherence, while achieving the best overall performance across most quantitative metrics. These results validate the proposed framework and highlight its potential for low-cost, highly controllable animation creation.

Meixi Li, Xianlin Zhang, Yue Zhang, Xueming Li• 2026

Related benchmarks

TaskDatasetResultRank
Animation SynthesisSAKUGA-42M (test)
EDMD4.1588
5
Sparse Key-Sketch Animation SynthesisAnimation 9 samples (test)
Overall Quality51.52
5
Showing 2 of 2 rows

Other info

Follow for update