Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Epipolar Geometry Improves Video Generation Models

About

Video generation models have advanced significantly through the latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts that break the illusion of realistic 3D scenes. 3D-consistent video generation could significantly impact numerous downstream applications in generation and reconstruction tasks. We explore how epipolar geometry constraints improve modern video diffusion models. Despite using massive training data, these models fail to capture fundamental geometric principles. We align diffusion models using pairwise epipolar geometry constraints via preference-based optimization, directly addressing unstable trajectories and geometric artifacts through mathematically principled geometric enforcement. Our approach efficiently enforces geometric principles without requiring end-to-end differentiability. Evaluation demonstrates that classical geometric constraints provide more stable optimization signals than modern learned metrics. Training on static scenes with dynamic cameras ensures metric quality while the model generalizes to various dynamic scenes. By bridging data-driven learning with classical computer vision, we reduce epipolar error by 31% and improve human-rated consistency from 54% to 72% without compromising visual quality.

Orest Kupyn, Th\'eo Uscidda, Marta Tintore Gazulla, Fabian Manhardt, Federico Tombari, Christian Rupprecht• 2025

Related benchmarks

TaskDatasetResultRank
Text-to-Video GenerationVBench--
209
Image-to-Video GenerationVBench
Motion Smoothness0.992
46
Video GenerationDL3DV RealEstate10K Static-scene benchmark
Epipolar Consistency0.098
10
Video GenerationStandard Video Evaluation Benchmark
Subject Consistency0.8407
10
Video GenerationMiraData Dynamic-scene benchmark
VQ Score45.5
8
Text-to-Video GenerationHuman evaluation--
8
Text-to-Video GenerationDL3DV-10K 1K (test)
PSNR17.69
7
Video GenerationVBench standard caption set
Subject Consistency96.01
5
Bidirectional Video GenerationGB3DV 25k
PSNR22.45
4
Image-to-Video GenerationDL3DV-10K 1K (test)
PSNR15.07
4
Showing 10 of 17 rows

Other info

Follow for update