Beyond IID: How General Are Tabular Foundation Models, Really?
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
arXiv:2606. 02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures.
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
arXiv:2609.13202v1 Announce Type: new Abstract: Feature engineering has long been a cornerstone of tabular machine learning. Tabular foundation models (TFMs) are pretrained on a wide range of tabular...
arXiv:2609.37989v1 Announce Type: new Abstract: Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, ta...
arXiv:2605. 28418v3 Announce Type: replace Abstract: With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult.
TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.
SymboLLM-FE combines symbolic regression and large language models to automate feature engineering for tabular data. It first extracts mathematically expressive formulas that correlate strongly with the target, then refines them with LLMs to improve interpretability. Experiments on six real‑world datasets and four Kaggle competitions show that SymboLLM‑FE outperforms existing AutoFE methods while reducing the number of costly LLM calls.
arXiv:2607. 23286v1 Announce Type: new Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space.
RelICL: Training-free Relational Learning with Tabular Foundation Models proposes a new method for relational learning that addresses two key issues of deep feature synthesis—feature explosion and interaction blindness—by propagating and fusing information step by step through the schema graph using a tabular foundation model. The approach retains the benefits of DFS while improving scalability and performance. Experiments on RelBench tasks show that RelICL performs on par with the strongest DFS-based approach.
arXiv:2603. 02221v2 Announce Type: replace-cross Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods.
Support-Compiled Feature Folding (SCFF) is a training‑free inference framework that addresses the feature‑side scaling dilemma in tabular foundation models by routing support‑ranked features through bounded leaves of the native encoder, checking residual evidence, and merging encoded messages for a single contextual prediction. This approach transforms quadratic pairwise mixing into linear‑in‑width work with a bounded local working set, achieving dataset‑macro accuracy and NLL improvements across six backbones on an 18‑dataset wide‑table slice. SCFF delivers significant GPU‑memory savings (median 2.09×–2.36×) and, when constrained by a peak‑memory ceiling, further boosts accuracy by up to 4.06 points over the widest single‑leaf baseline. whyItMatters":"SCFF demonstrates that memory‑efficient inference can simultaneously improve accuracy and reduce resource usage in tabular foundation models, offering a practical solution for deploying these models at scale."
The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.
arXiv:2608. 04174v1 Announce Type: new Abstract: Time series data are ubiquitous in practical applications, where classification (TSC) and extrinsic regression (TSER) have emerged as essential tasks for obtaining value from temporal sequences.