arXiv:2609.37959v1 Announce Type: new
Abstract: Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We p...
By Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
The paper investigates how Tabular Foundation Models (TFMs) can achieve strong transfer learning by self‑supervised pre‑training on a single real table. It finds that a table’s usefulness is largely determined by its number of features rather than instances, and that fine‑grained column‑level preprocessing improves downstream performance while dataset‑level filtering does not. The authors propose that tabular in‑context generalization is primarily retrieval‑based, with models learning to identify and aggregate relevant examples from the provided context.
By Nour Shaheen, Junwei Ma, Alex Labach, Frank Hutter, Valentin Thomas, Anthony L. Caterini
The paper investigates how Tabular Foundation Models (TFMs) can achieve strong transfer learning by self‑supervised pre‑training on a single real table, rather than large synthetic or real datasets. It finds that a table’s usefulness for downstream tasks is mainly determined by the number of features, not instances, and that fine‑grained column‑level preprocessing improves performance while dataset‑level filtering does not. The authors propose a task‑centric, retrieval‑based view of in‑context generalization, suggesting that effective TFMs identify and aggregate relevant examples from the provided context.
arXiv:2606. 11640v1 Announce Type: cross Abstract: Few-shot tabular learning provides a cost-effective approach for real-world applications where annotation is costly and collecting sufficient samples for new tasks is difficult.
By Ruxue Shi, Yili Wang, Mengnan Du, Hangting Ye, Yi Chang, Xin Wang
The paper critically evaluates common few‑shot learning protocols that rely on pre‑training a model on a large auxiliary set with classes disjoint from the target but drawn from the same visual domain. By comparing no pre‑training, class‑disjoint in‑domain pre‑training, supervised out‑of‑domain pre‑training, and label‑free out‑of‑domain pre‑training across eight datasets and three architectures, the authors find that in‑domain pre‑training yields a 33.41‑point average improvement, while out‑of‑domain pre‑training offers a 23.75‑point gain, revealing a 9.66‑point optimistic bias due to domain overlap. They also demonstrate that a label‑free augmentation strategy can match supervised out‑of‑domain performance and propose a descriptor‑based source‑selection method that closely approximates oracle selection, underscoring the need to move beyond in‑domain pre‑training as the default evaluation protocol.
By Alejandro Galan-Cuenca, Marcelo Saval-Calvo, Antonio Javier Gallego
The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.
By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas