VINO: A Unified Visual Generator with Interleaved OmniModal Context

About

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion backbone that conditions on text, images and videos, enabling a broad range of visual creation and editing tasks under one model. Specifically, VINO couples a vision-language model (VLM) with a Multimodal Diffusion Transformer (MMDiT), where multimodal inputs are encoded as interleaved conditioning tokens, and then used to guide the diffusion process. This design supports multi-reference grounding, long-form instruction following, and coherent identity preservation across static and dynamic content, while avoiding modality-specific architectural components. To train such a unified system, we introduce a multi-stage training pipeline that progressively expands a video generation base model into a unified, multi-task generator capable of both image and video input and output. Across diverse generation and editing benchmarks, VINO demonstrates strong visual quality, faithful instruction following, improved reference and attribute preservation, and more controllable multi-identity edits. Our results highlight a practical path toward scalable unified visual generation, and the promise of interleaved, in-context computation as a foundation for general-purpose visual creation.

Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye• 2026

Related benchmarks

Task	Dataset	Result
Text-to-Image Generation	GenEval (test)	Two Obj. Acc88	250
Text-to-Video Generation	VBench	Quality Score84	209
Video Understanding	Video-MME without subtitles	Overall Score69.3	145
Vision Understanding	MMMU	--	71
Video Editing	OpenVE-Bench	Overall Score3.11	39
subject-to-video generation	OpenS2V	Total59.31	32
Instruction-Guided Video Editing	OpenVE-Bench	Overall Score4.34	29
Video Editing	OpenVE-Bench (test)	Overall Score4.34	28
Subject-to-video	OpenS2V Eval	Total Score57.85	23
Compositional Multi-Image-to-Video Generation	IntelligentVBench 2Subjects with BKG	IF Score3.56	21

Showing 10 of 36 rows

Other info

GitHub

Follow for update

@wizwand_team Discord