arXiv Machine Learning By Andrej Tschalzev, Nick Erickson, Yuyang Wang, Huzefa Rangwala, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

Read the original on arXiv Machine Learning →

arXiv:2606. 02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 30

Beyond IID: How General Are Tabular Foundation Models, Really?

arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.

By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv Machine Learning
Sep 17

TabICLv2: A better, faster, scalable, and open tabular foundation model

TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.

By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
arXiv Machine Learning
Aug 31

SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data

SymboLLM-FE combines symbolic regression and large language models to automate feature engineering for tabular data. It first extracts mathematically expressive formulas that correlate strongly with the target, then refines them with LLMs to improve interpretability. Experiments on six real‑world datasets and four Kaggle competitions show that SymboLLM‑FE outperforms existing AutoFE methods while reducing the number of costly LLM calls.

By Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo