Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation

About

We introduce Quantized Language-Image Pretraining (QLIP), a visual tokenization method that combines state-of-the-art reconstruction quality with state-of-the-art zero-shot image understanding. QLIP trains a binary-spherical-quantization-based autoencoder with reconstruction and language-image alignment objectives. We are the first to show that the two objectives do not need to be at odds. We balance the two loss terms dynamically during training and show that a two-stage training pipeline effectively mixes the large-batch requirements of image-language pre-training with the memory bottleneck imposed by the reconstruction objective. We validate the effectiveness of QLIP for multimodal understanding and text-conditioned image generation with a single model. Specifically, QLIP serves as a drop-in replacement for the visual encoder for LLaVA and the image tokenizer for LlamaGen with comparable or even better performance. Finally, we demonstrate that QLIP enables a unified mixed-modality auto-regressive model for understanding and generation.

Yue Zhao, Fuzhao Xue, Scott Reed, Linxi Fan, Yuke Zhu, Jan Kautz, Zhiding Yu, Philipp Kr\"ahenb\"uhl, De-An Huang• 2025

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE--
2056
Visual Question AnsweringTextVQA--
1455
Visual Question AnsweringGQA
Accuracy61.8
1445
Visual Question AnsweringChartQA
Accuracy14.1
620
Multimodal UnderstandingMMStar
Accuracy39.9
511
Diagram Question AnsweringAI2D
AI2D Accuracy61.9
509
Optical Character RecognitionOCRBench
Score290
486
Multi-discipline Multimodal UnderstandingMMMU--
422
Diagram UnderstandingAI2D
Accuracy61.9
377
Visual Question AnsweringInfoVQA
Accuracy14.8
264
Showing 10 of 29 rows

Other info

Follow for update