Support-Compiled Feature Folding (SCFF) is a training‑free inference framework that addresses the feature‑side scaling dilemma in tabular foundation models by routing support‑ranked features through bounded leaves of the native encoder, checking residual evidence, and merging encoded messages for a single contextual prediction. This approach transforms quadratic pairwise mixing into linear‑in‑width work with a bounded local working set, achieving dataset‑macro accuracy and NLL improvements across six backbones on an 18‑dataset wide‑table slice. SCFF delivers significant GPU‑memory savings (median 2.09×–2.36×) and, when constrained by a peak‑memory ceiling, further boosts accuracy by up to 4.06 points over the widest single‑leaf baseline.
whyItMatters":"SCFF demonstrates that memory‑efficient inference can simultaneously improve accuracy and reduce resource usage in tabular foundation models, offering a practical solution for deploying these models at scale."
By Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
arXiv:2605.12904v2 Announce Type: replace
Abstract: Tabular foundation models (TFMs) have emerged as a powerful paradigm for in-context learning on structured data, enabling direct prediction on new...
By Yilong Chen, Xueying Ding, Leman Akoglu
arXiv:2605. 28418v3 Announce Type: replace Abstract: With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult.
By Markus Herre, Andrej Tschalzev, Sascha Marton, Christian Bartelt
TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.
By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
arXiv:2609.37989v1 Announce Type: new
Abstract: Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, ta...
By Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan, Sujay Sanghavi, Abhimanyu Das, Weihao Kong
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv:2607. 11007v1 Announce Type: new Abstract: Few-shot multimodal classification commonly attaches a lightweight head, such as $k$-nearest neighbors, logistic regression, or a linear SVM, to a frozen pretrained encoder.
By Jingxiang Zhang, Lujia Zhong, Zijie Zhu, Shuo Huang, Yuang Xu
arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.
By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati
The paper investigates how synthetic pretraining priors used in tabular foundation models (TFMs) influence downstream performance. By reconstructing the synthetic data generators of four TFMs and comparing their generated tasks to two popular tabular benchmarks using structural descriptors, the authors measure structural coverage and normalized density. They find that some generators provide broader and denser support for benchmark tasks, and that stronger synthetic-to-benchmark support generally correlates with better model performance.
By He Zhao, Ryan Thompson, Daniel M. Steinberg, Ashfaqur Rahman, Edwin V. Bonilla, Cheng Soon Ong
The paper introduces RefineICL, an attention‑gated, feed‑forward‑network‑free framework that refines representations in situ for tabular foundation models. By using support labels to guide episode‑specific updates, the method transfers learned corrections to unlabeled queries without altering model parameters, achieving state‑of‑the‑art performance on AMLB29 and TabArena benchmarks. Experiments and internal interventions demonstrate that intermediate support updates are essential for constructing task‑specific predictors in context.
By Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
Xiaomi-TabLDM is a tabular foundation model that performs classification and regression via in-context learning without task‑specific fine‑tuning. It is pretrained solely on synthetic data from structural causal models, achieving top‑ranked regression results on multiple benchmarks while reducing training and prediction time compared to leading models. The architecture incorporates a three‑stage training strategy, dual‑stream feature grouping, lightweight attention residuals, and sparse mixture‑of‑experts, and it can further improve accuracy through test‑time compute scaling.
By TabLDM Team, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang, Chunxiao Liu, Erli Meng, Bin Wang
arXiv:2608. 01400v1 Announce Type: new Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity.
By Rasa Hosseinzadeh, Alex Labach, Zexin Xue, Shuyi Han, Valentin Thomas, Anthony L. Caterini