Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

About

High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in $\mathbb{R}^m$ ill-conditioned since $n \ll m$. We propose BSTabDiff, a block-subunit generative framework that partitions the $m$ observed features into $M$ latent blocks ($M \ll m$) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space $\mathbb{R}^M$ while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.

Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh• 2026

Related benchmarks

TaskDatasetResultRank
Tabular SynthesisCol
Accuracy83.26
11
Tabular SynthesisLNG
Accuracy95.96
11
Tabular SynthesisGLI
Accuracy82.35
11
Tabular SynthesisSMK
Accuracy72.08
11
Tabular SynthesisAML
Accuracy95.3
11
Tabular SynthesisPRS
Accuracy90.76
11
Tabular SynthesisARC
Accuracy86.5
11
Tabular SynthesisTox
Accuracy87.7
11
ClassificationColon (COL) (downstream performance folds)
TANDEM Accuracy81.97
2
Showing 9 of 9 rows

Other info

Follow for update