Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models

About

Evaluating video generation with clean, pixel-based reward models disconnects evaluation from the noisy diffusion process and incurs massive VAE decoding costs. In this paper, we challenge this paradigm by asking a fundamental question: Can a powerful video generator inherently discriminate preferences directly from noisy latents? To answer this, we introduce \textbf{PRISM} (\textbf{P}reference \textbf{R}epresentation in \textbf{I}ntermediate \textbf{S}tates of Diffusion \textbf{M}odels). PRISM employs a lightweight Query-based Aggregation head with a frozen video diffusion backbone to decode preference signals from noisy latents. Surprisingly, PRISM not only achieves SOTA preference accuracy but also unlocks strong noise-robustness, which enables early-stage Best-of-$N$ sampling. This allows for filtering suboptimal candidates at the very beginning of denoising, drastically reducing computation while boosting video quality. We also reveal a strong positive correlation between a backbone's generative performance and its inherent evaluative power, enabling self-improving video backbones.

Haoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li, Hongzheng Yang, Wei Liu• 2026

Related benchmarks

TaskDatasetResultRank
Video Generation EvaluationVBench
Quality Score86.3384
25
Preference PredictionVideoGen-RewardBench
Score @ Threshold 9976.81
12
Preference PredictionVLRM-Bench
Preference Prediction Score (k=99)80.25
12
Showing 3 of 3 rows

Other info

Follow for update