Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images

About

We present a hierarchical VAE that, for the first time, generates samples quickly while outperforming the PixelCNN in log-likelihood on all natural image benchmarks. We begin by observing that, in theory, VAEs can actually represent autoregressive models, as well as faster, better models if they exist, when made sufficiently deep. Despite this, autoregressive models have historically outperformed VAEs in log-likelihood. We test if insufficient depth explains why by scaling a VAE to greater stochastic depth than previously explored and evaluating it CIFAR-10, ImageNet, and FFHQ. In comparison to the PixelCNN, these very deep VAEs achieve higher likelihoods, use fewer parameters, generate samples thousands of times faster, and are more easily applied to high-resolution images. Qualitative studies suggest this is because the VAE learns efficient hierarchical visual representations. We release our source code and models at https://github.com/openai/vdvae.

Rewon Child• 2020

Related benchmarks

Task	Dataset	Result
Image Generation	CIFAR-10 (test)	--	536
Density Estimation	CIFAR-10 (test)	Bits/dim2.87	134
Image Generation	FFHQ	FID33.5	91
Density Estimation	ImageNet 64x64 (test)	Bits Per Sub-Pixel3.52	71
Density Estimation	ImageNet 32x32 (test)	Bits per Sub-pixel3.8	69
Generative Modeling	CIFAR-10 (test)	NLL (bits/dim)2.87	62
Density Estimation	CIFAR-10	bpd2.87	40
Unconditional image synthesis	FFHQ 256x256 (test)	FID28.5	31
Unconditional Image Generation	FFHQ 256x256 (test)	FID28.5	25
Unconditional image modeling	ImageNet 64x64	Bits/Dim3.52	17

Showing 10 of 26 rows

Other info

Code

Follow for update

@wizwand_team Discord