arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.
By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.
By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas
TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.
By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
The paper introduces SynPro, a synthetic data generation framework that augments limited organic text for large language model pretraining by applying rephrasing and reformatting operations. SynPro’s generators are optimized with reinforcement learning rewards for quality, faithfulness, and data influence, and are updated continuously as training plateaus. Experiments on 400M, 1.1B, and 2B models show that SynPro can unlock 3.4–5.2× the effective tokens of standard repetition, even outperforming a non‑data‑bound oracle at larger scales.
By Zichun Yu, Chenyan Xiong
arXiv:2406. 08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate causal relationships that go beyond pairwise correlations largely unmeasured.
By Zineb Senane, Axel Karlsson, Lele Cao, Oleg Smirnov, Cheng Zhang, Sahar Asadi, Hedvig Kjellstr\"om, Gustav Eje Henter, Ruibo Tu
The paper investigates how synthetic pretraining priors used in tabular foundation models (TFMs) influence downstream performance. By reconstructing the synthetic data generators of four TFMs and comparing their generated tasks to two popular tabular benchmarks using structural descriptors, the authors measure structural coverage and normalized density. They find that some generators provide broader and denser support for benchmark tasks, and that stronger synthetic-to-benchmark support generally correlates with better model performance.
By He Zhao, Ryan Thompson, Daniel M. Steinberg, Ashfaqur Rahman, Edwin V. Bonilla, Cheng Soon Ong
arXiv:2412. 06265v3 Announce Type: replace Abstract: Deep tabular models should ideally balance predictive performance, parameter efficiency, and robustness to imperfect learning signals---properties that are rarely considered jointly.
By Seungeun Lee, Kihwan Lee, Subin Bae, Sangjun Lee, Seulbin Lee, Julia Stoyanovich, Il-Youp Kwak, Seungsang Oh
arXiv:2606. 02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures.
By Andrej Tschalzev, Nick Erickson, Yuyang Wang, Huzefa Rangwala, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
The paper introduces In-Table Prediction (ITB), a self‑supervised task where deep neural networks learn to predict any column in a table from the remaining columns. It proposes a novel neural layer to handle missing continuous values, generates synthetic datasets with controlled column relationships, and evaluates three architectures—MLP, ResNet, and Transformer—showing that attention‑based Transformers perform best when ample training data and large embeddings are used. The study is limited to synthetic, small‑column tables and is presented as an initial investigation rather than a comprehensive real‑world analysis.
By Xiao Zhao, Daniela Oelke
arXiv:2605. 28198v2 Announce Type: replace Abstract: Existing approaches for synthetic tabular data generation are based on either purely generative models or LLMs, both of which struggle with data heterogeneity, logical consistency, rare-event coverage, and robustness in low-data regimes.
By Junfeng Nie, Alvin Jin, Xiaohui Chen
arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.
By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati