arXiv Machine Learning

Parameter-Free Encoders Remain Viable for RDB Foundation Models

arXiv:2607. 05476v1 Announce Type: new Abstract: Given a relational database (RDB) storing heterogeneous tabular information, how can we predict missing (or future) values in some target column of interest?

arXiv AI
Sep 15

Generalization Can Emerge in Tabular Foundation Models From a Single Table

The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.

By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas
Hugging Face Trending Papers
Jun 29

FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks

We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders. Unlike existing tabular in-context learners, which entangle feature representations with a specific prediction target, our design produces \textit{target-agnostic} row embeddings that can be leveraged across a wide range of downstream tasks within a table-native in-context learning setup.

arXiv Machine Learning
1d ago

STEER: Reducing Inference Cost in Relational Foundation Models through Semantically Informed Sampling

STEER is a sampling method for relational foundation models that reduces inference cost by focusing on the most relevant tables for a prediction task. It uses a large language model to rank foreign‑key edges in the database schema into relevance tiers, then assigns traversal probabilities based on these tiers. Evaluated on three state‑of‑the‑art RFMs, STEER cuts inference context size by roughly 40% on average while preserving or improving accuracy.

By Abdalla Mohamed, Ashraf Aboulnaga
arXiv Machine Learning
Sep 2

Can LLMs Use Relational Transformer Embeddings?

The paper investigates whether large language models (LLMs) can leverage frozen relational‑transformer embeddings by injecting them as soft tokens. Using a learned MLP projection and LoRA adaptation, the authors fine‑tune Qwen3.5‑4B on chain‑of‑thought reasoning traces and group‑based reinforcement learning, then evaluate on ten binary classification tasks across six RelBench databases. The hybrid approach consistently underperforms the standalone relational transformer, showing sensitivity to serialization format, token budget, and RL stability, leading the authors to conclude that stronger alignment objectives and schema‑aware design are needed for reliable relational prediction.

By Francisco Galuppo Azevedo, Clarissa Lima Loures