arXiv Machine Learning

Memory Efficient Tabular Foundation Models

arXiv:2607. 27546v1 Announce Type: new Abstract: Tabular Foundation Models, such as TabPFN, have received a large amount of recent attention due to their performance on in-context tabular machine learning tasks, which often exceeds classical baselines.

arXiv Machine Learning
Aug 24

Tydra: An Efficient Hybrid Model for Tabular Data

Tydra is a hybrid Transformer‑State Space Model that interleaves attention and SSM layers for tabular in‑context learning. It achieves a 30% reduction in inference time compared to the Transformer‑only TabPFN while preserving most of its predictive performance. On 30 OpenML datasets, Tydra also outperforms a Hydra model that is roughly ten times larger, demonstrating that hybrid architectures can balance accuracy and efficiency for tabular foundation models.

By Mieszko Komisarczyk, Saurabh Mathur, Maurice Kraus, Sriraam Natarajan, Kristian Kersting
arXiv Machine Learning
Jun 4

Towards Pretraining Text Encoders for TabPFN

arXiv:2606. 04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handle high-cardinality text features.

By Mustafa Tajjar, Alexander Pfefferle, Lennart Purucker, Frank Hutter
arXiv Machine Learning
Sep 17

TabICLv2: A better, faster, scalable, and open tabular foundation model

TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.

By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
arXiv Machine Learning
Sep 24

Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models

Support-Compiled Feature Folding (SCFF) is a training‑free inference framework that addresses the feature‑side scaling dilemma in tabular foundation models by routing support‑ranked features through bounded leaves of the native encoder, checking residual evidence, and merging encoded messages for a single contextual prediction. This approach transforms quadratic pairwise mixing into linear‑in‑width work with a bounded local working set, achieving dataset‑macro accuracy and NLL improvements across six backbones on an 18‑dataset wide‑table slice. SCFF delivers significant GPU‑memory savings (median 2.09×–2.36×) and, when constrained by a peak‑memory ceiling, further boosts accuracy by up to 4.06 points over the widest single‑leaf baseline. whyItMatters":"SCFF demonstrates that memory‑efficient inference can simultaneously improve accuracy and reduce resource usage in tabular foundation models, offering a practical solution for deploying these models at scale."

By Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun