Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Shift-and-Sum Quantization for Visual Autoregressive Models

About

Post-training quantization (PTQ) enables efficient deployment of deep networks using a small set of data. Its application to visual autoregressive models (VAR), however, remains relatively unexplored. We identify two key challenges for applying PTQ to VAR: (i) large reconstruction errors in attention-value products, especially at coarse scales where high attention scores occur more frequently; and (ii) a discrepancy between the sampling frequencies of codebook entries and their predicted probabilities due to limited calibration data. To address these challenges, we propose a PTQ framework tailored for VAR. First, we introduce a shift-and-sum quantization method that reduces reconstruction errors by aggregating quantized results from symmetrically shifted duplicates of value tokens. Second, we present a resampling strategy for calibration data that aligns sampling frequencies of codebook entries with their predicted probabilities. Experiments on class-conditional image generation, inpainting, outpainting, and class-conditional editing show consistent improvements across VAR architectures, establishing a new state of the art in PTQ for VAR.

Jaehyeon Moon, Bumsub Ham• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Image GenerationGenEval
GenEval Score72.5
459
Text-to-Image GenerationHPS v2.1
Overall Score32.08
153
Semantic segmentationCOCO
mIoU67
119
Text-to-Image GenerationImageReward
ImageReward Score0.929
119
Class-conditional Image GenerationImageNet 2009 (val)
Inception Score (IS)269.9
52
Panoptic SegmentationCOCO
PQ57.8
46
Instance SegmentationCOCO
AP0.486
21
Class-conditional EditingVAR-d30
CLIP Score24.65
9
Image InpaintingVAR-d30
LPIPS0.2835
9
Image OutpaintingVAR-d30
LPIPS0.2134
9
Showing 10 of 10 rows

Other info

Follow for update