[Paper Review] Why Tabular Foundation Models Should Be a Research Priority
This position paper argues for prioritizing research into Large Tabular Models (LTMs), a new class of foundation models for tabular data, which could revolutionize data science by enabling contextualized, few-shot, and synthetic data generation across domains. The authors advocate for LTMs as a high-impact, underexplored frontier with strong potential for scientific discovery, privacy-preserving data sharing, and robust, inclusive machine learning.
Recent text and image foundation models are incredibly impressive, and these models are attracting an ever-increasing portion of research resources. In this position piece we aim to shift the ML research community's priorities ever so slightly to a different modality: tabular data. Tabular data is the dominant modality in many fields, yet it is given hardly any research attention and significantly lags behind in terms of scale and power. We believe the time is now to start developing tabular foundation models, or what we coin a Large Tabular Model (LTM). LTMs could revolutionise the way science and ML use tabular data: not as single datasets that are analyzed in a vacuum, but contextualized with respect to related datasets. The potential impact is far-reaching: from few-shot tabular models to automating data science; from out-of-distribution synthetic data to empowering multidisciplinary scientific discovery. We intend to excite reflections on the modalities we study, and convince some researchers to study large tabular models.
Motivation & Objective
- To highlight the underrepresentation of tabular data in foundation model research despite its dominance in real-world applications.
- To argue that tabular foundation models (LTMs) present a high-impact, underexplored frontier comparable to text and vision foundation models.
- To address the reasons why LTMs have been overlooked: lack of large-scale tabular metadatasets, inherent difficulty in tabular ML, and human perception biases favoring vision and text.
- To propose that LTMs can enable transformative applications such as few-shot data augmentation, synthetic data generation, and cross-domain data integration.
- To encourage researchers to redirect attention toward LTMs, emphasizing their potential for public good, inclusivity, and lower barriers to entry compared to large language models.
Proposed method
- Propose the concept of Large Tabular Models (LTMs) as foundation models tailored for tabular data, analogous to LLMs and vision FMs.
- Frame LTMs as capable of generating synthetic tabular data, performing cross-dataset joins, and enriching datasets with missing columns via few-shot learning.
- Suggest using LTM embeddings for downstream tasks, such as model fine-tuning or integration with other foundation models.
- Emphasize the need for large-scale, diverse, and high-quality tabular metadatasets to train LTMs, similar to how ImageNet enabled vision FMs.
- Advocate for evaluation protocols that assess generalization, bias, and reliability across diverse domains and data distributions.
- Highlight the importance of traceability and identity verification in tabular data to reduce misuse risks compared to generative image or text models.

Experimental results
Research questions
- RQ1Why has tabular data, despite its prevalence in science and industry, received minimal attention in foundation model research compared to text and vision?
- RQ2What unique challenges do tabular data pose for foundation model development, and how can they be addressed?
- RQ3How can Large Tabular Models (LTMs) enable few-shot data augmentation, synthetic data generation, and cross-dataset reasoning?
- RQ4What are the potential impacts of LTMs on data democratization, privacy, scientific discovery, and model robustness?
- RQ5How can LTMs be evaluated reliably, and what safeguards are needed to prevent bias and misuse?
Key findings
- Tabular data is the dominant data modality in science, healthcare, finance, and government, yet it is vastly under-represented in foundation model research.
- Despite strong baselines like XGBoost, tabular ML has not advanced as rapidly as vision or NLP, leaving significant room for innovation in foundation models.
- LTMs can generate high-quality synthetic tabular data that preserves statistical properties and enables privacy-preserving data sharing.
- LTMs can perform cross-dataset reasoning, such as inferring missing columns or joining datasets across domains using few-shot prompts.
- The risk of misuse for LTMs is lower than for text and vision FMs due to data traceability and the difficulty of fabricating credible tabular data without domain expertise.
- LTM research offers a high-impact, accessible frontier with lower computational barriers than LLMs, enabling broader participation and faster progress.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.