arXiv:2607. 05476v1 Announce Type: new Abstract: Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of interest?
By Linjie Xu, David Wipf
arXiv:2607. 29129v1 Announce Type: new Abstract: Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline.
By Mohammad Sadeq Abolhasani, Viswanath Ganapathy
arXiv:2606. 30336v1 Announce Type: new Abstract: We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders.
By Marek Polewczyk, Maximilian Schambach, Marco Spinaci, Sam Thelin, Johannes H\"ohne
arXiv:2608. 10837v1 Announce Type: cross Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs.
By Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger
STEER is a sampling method for relational foundation models that reduces inference cost by focusing on the most relevant tables for a prediction task. It uses a large language model to rank foreign‑key edges in the database schema into relevance tiers, then assigns traversal probabilities based on these tiers. Evaluated on three state‑of‑the‑art RFMs, STEER cuts inference context size by roughly 40% on average while preserving or improving accuracy.
By Abdalla Mohamed, Ashraf Aboulnaga
The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.
By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas