Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

About

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinders truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. UniAR adapts a pretrained vision encoder with multi-level feature fusion and a lookup-free bitwise quantization scheme, preserving both high-level semantics and low-level details while scaling the effective visual vocabulary at minimal cost. Building on this, the unified autoregressive model adopts parallel-bitwise-prediction to jointly predict spatially grouped, multi-level visual codes, substantially reducing visual sequence length and accelerating generation. Finally, a diffusion-based visual decoder operates on discrete visual tokens to decode high-fidelity images. Through large-scale pre-training, followed by supervised fine-tuning and reinforcement learning, UniAR achieves state-of-the-art performance on image generation and image editing while remaining competitive on multimodal understanding benchmarks. The project page is available at https://sharelab-sii.github.io/uniar-web.

Wujian Peng, Lingchen Meng, Yuxuan Cai, Xianwei Zhuang, Yuhuan Yang, Rongyao Fang, Chenfei Wu, Junyang Lin, Zuxuan Wu, Shuai Bai• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Image GenerationGenEval
Overall Score0.86
318
Image EditingImgEdit-Bench
Overall Score3.73
256
Multi-modal Video UnderstandingMVBench
Score62.3
84
Text RenderingLongText-Bench English
Score0.917
39
Multimodal Image UnderstandingDocVQA
Score91.4
13
Multimodal Image UnderstandingInfoVQA
Score70
12
Text RenderingOneIG-Bench English
Score0.873
12
Multimodal Image UnderstandingOCRBench
Score83.3
11
Multimodal UnderstandingChartQA
Score75.9
4
Multimodal UnderstandingRLWDQA
Score68.5
3
Showing 10 of 10 rows

Other info

GitHub

Follow for update