The paper investigates whether training classical machine learning models remains worthwhile when large language models (LLMs) can label tabular data without training. By defining a labeled‑data crossover point (N*) where a trained classical model surpasses a frozen LLM’s flat error, the authors analyze 126 student evaluations of GPT models across 18 datasets and compare them to power‑law learning curves of six classical model families. Results show that in 86% of cases a classical model outperforms the LLM with no more labeled data than already available, and the crossover occurs at a median of about 6% of the training set, suggesting that collecting a few hundred labels and training a gradient‑boosted model is typically advantageous.
By Kaihua Ding
arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.
By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv:2609.16309v1 Announce Type: new
Abstract: Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains...
By Renat Sergazinov, Artem Chistyakov, Sergey Pankevich, Artem Babenko
arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.
By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
Support-Compiled Feature Folding (SCFF) is a training‑free inference framework that addresses the feature‑side scaling dilemma in tabular foundation models by routing support‑ranked features through bounded leaves of the native encoder, checking residual evidence, and merging encoded messages for a single contextual prediction. This approach transforms quadratic pairwise mixing into linear‑in‑width work with a bounded local working set, achieving dataset‑macro accuracy and NLL improvements across six backbones on an 18‑dataset wide‑table slice. SCFF delivers significant GPU‑memory savings (median 2.09×–2.36×) and, when constrained by a peak‑memory ceiling, further boosts accuracy by up to 4.06 points over the widest single‑leaf baseline.
whyItMatters":"SCFF demonstrates that memory‑efficient inference can simultaneously improve accuracy and reduce resource usage in tabular foundation models, offering a practical solution for deploying these models at scale."
By Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
arXiv:2608. 05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap.
By Xinran Feng, Yi Xie, Chao Zhang, Ruikun Li, Wanyun Ling, Ziyue Li, Chenxi Liu
arXiv:2609.37076v1 Announce Type: new
Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
By Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen
arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.
By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati
arXiv:2409. 02228v2 Announce Type: replace Abstract: When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change?
By Eric Zhang, Leshem Choshen, Jacob Andreas
arXiv:2609.16454v1 Announce Type: new
Abstract: Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend...
By Kirill Skobelev, Eric Fithian, X. Y. Han
arXiv:2609.37989v1 Announce Type: new
Abstract: Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, ta...
By Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan, Sujay Sanghavi, Abhimanyu Das, Weihao Kong
arXiv:2609.37891v1 Announce Type: cross
Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko