Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

About

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-end over fixed-length sequences and cannot scale, generate shot-by-shot with memory banks that grow linearly, or orchestrate pretrained generators under an LLM planner without a multi-shot-aware backbone. We present UnityShots, a memory-driven multi-shot audio-video generation system built on LTX-2.3, trained on annotated cinematic and music-video shots. The video stream maintains two fixed-size slots, a long-term memory (LTM) slot anchored to the opening shot and a short-term memory (STM) slot holding the immediately preceding tail, both updated at every cut by a boundary-conditioned gate that fuses visual cut probability and beat-tracker signals. The audio stream injects a reference speaker token at every shot to preserve vocal timbre without a sliding audio bank. A discrete cut-type prior, learned through AdaLN, becomes an inference-time control knob over transition strength. We release a benchmark of $200$ multi-cultural multi-shot sequences spanning six ethnic regions and ten or more languages, with per-shot reference identities, reference audio, and per-boundary transition labels. Evaluated across I2V, T2V, and R2V conditioning modes, UnityShots leads open-source baselines on every cross-shot coherence metric and matches the strongest closed-source system on the multi-shot axes.

Jiehui Huang, Yuechen Zhang, Bin Xia, Jiahao Wang, Xu He, Zhenchao Tang, Meng Chu, Xin Tao, Pengfei Wan, Jiaya Jia• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Video Generation200-sequence multi-cultural benchmark
NC Score4.82
10
Multi-shot Generationmulti-cultural benchmark I2V
Temporal Alignment (TA)20.62
6
Multi-shot Video GenerationMulti-cultural benchmark 3-4 shots
NC Score4.98
4
Multi-shot Video GenerationMulti-cultural benchmark 5 shots
NC Score4.91
4
Multi-shot Video GenerationMulti-cultural benchmark 6+ shots
NC Score5
4
Multi-shot Generationmulti-cultural benchmark T2V
Text Alignment Score (TA)19.17
2
Multi-shot Generationmulti-cultural benchmark (R2V)
Temporal Alignment (TA)17.98
2
Showing 7 of 7 rows

Other info

GitHub

Follow for update