Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Missing Data Imputation under Manifold Hypothesis

About

The manifold hypothesis posits that high-dimensional data are concentrated near a low-dimensional embedded manifold. Recent advances in mixture variational autoencoders (VAEs) provide a powerful tool for extracting such underlying structure in a faithful manner. The resulting geometric structure naturally introduces local and global relationships among variables, thereby providing a systematic way of imputing missing data. We propose a model-based imputation method that enables sampling from \( p(\bm{x}_{\mathrm{mis}} \mid \bm{x}_{\mathrm{obs}}) \) via a sampling-importance-resampling (SIR) procedure, which can be further augmented with a joint diffusion model in the latent space. Our method imputes missing data while respecting the underlying geometry, achieves competitive performance compared to state-of-the-art procedures, quantifies uncertainty in the imputations, and is model-based, thereby enabling on-the-fly imputation without rerunning the entire procedure.

Zelong Bi, Amuchechukwu Ibenegbu• 2026

Related benchmarks

TaskDatasetResultRank
RegressionSuperconduct
RMSE0.2661
228
Data ImputationSuperconductivity
Wasserstein Distance0.6334
216
Data ImputationWINE (test)
RMSE1.0436
205
Imputationfacialexpression
RMSE0.0667
195
ImputationSuperconductivity (MCAR)
RMSE0.2153
72
ImputationSuperconductivity (MAR)
RMSE0.2154
72
ImputationSuperconductivity MNAR
RMSE0.2724
70
Imputationpowerplant MAR
RMSE0.8511
67
Data Imputationpowerplant MAR (test)
Wasserstein Distance0.3143
66
ImputationFacial expression dataset (MNAR)
Wasserstein Distance0.5063
66
Showing 10 of 36 rows

Other info

Follow for update