arXiv Machine Learning By Yijun Lin, Sai Li

Context-Constrained Transfer Learning for Tabular Foundation Models via Data Distillation

Read the original on arXiv Machine Learning →

arXiv:2607. 04809v1 Announce Type: cross Abstract: Tabular Foundation Models (TFMs) have demonstrated strong empirical performance as black-box inference engines through in-context learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 31

Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data

The paper introduces CATTLE, a transfer learning framework for disjoint tabular datasets that eliminates the need for shared features by leveraging generalized context learned through transformer projection weights. By using key, value, and query weights from source and target domains, CATTLE performs cross‑domain attention transfer in a data‑agnostic manner. Experiments on ten source‑target pairs demonstrate that CATTLE outperforms nine state‑of‑the‑art baselines, achieving the best average rank (2.9) and a 3.7% AUROC improvement.

By Kazi F. Akhter, Ibna Kowsar, Manar D. Samad
arXiv Machine Learning
23h ago

Distillation of Tabular Foundation Models into Efficient Predictors

The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.

By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv AI
Sep 15

Generalization Can Emerge in Tabular Foundation Models From a Single Table

The paper demonstrates that a tabular foundation model can achieve strong generalization using only a single real table for self‑supervised pre‑training, challenging the belief that large synthetic or real datasets are necessary. By systematically pre‑training and evaluating across diverse benchmarks, the authors show that the number and quality of tasks that can be derived from a dataset are critical for downstream performance. This finding suggests that carefully constructed task sets from limited data can enable effective transfer learning in tabular models.

By Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi, Frank Hutter, Anthony L. Caterini, Valentin Thomas