Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
About
Foundation models for tabular data, like TabPFN, achieve strong performance on small datasets when pre-trained solely on synthetic data. We show that this performance can be significantly boosted by a targeted continued pre-training phase. Specifically, we demonstrate that leveraging a small, curated collection of large, real-world datasets for continued pre-training yields superior downstream predictive accuracy compared to using broader, potentially noisier corpora like CommonCrawl or GitTables. Our resulting model, Real-TabPFN, achieves substantial performance gains on 29 datasets from the OpenML AutoML Benchmark.
Anurag Garg, Muhammad Ali, Noah Hollmann, Lennart Purucker, Samuel M\"uller, Frank Hutter• 2025
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Classification | BCCO-CLS | AUC85.27 | 12 | |
| Classification | GI-CLS | AUC0.8967 | 9 | |
| Classification | Talent CLS | AUC89.68 | 9 | |
| Classification | Tabarena CLS | AUC0.8525 | 9 | |
| Classification | Tabzilla CLS | AUC90.96 | 9 |
Showing 5 of 5 rows