arXiv AI By Shu Wan, Abhinav Gorantla, Huan Liu, K. Sel\c{c}uk Candan

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Read the original on arXiv AI →

arXiv:2605. 29411v2 Announce Type: replace-cross Abstract: Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

Xiaomi-TabLDM is a tabular foundation model that performs classification and regression via in-context learning without task‑specific fine‑tuning. It is pretrained solely on synthetic data from structural causal models, achieving top‑ranked regression results on multiple benchmarks while reducing training and prediction time compared to leading models. The architecture incorporates a three‑stage training strategy, dual‑stream feature grouping, lightweight attention residuals, and sparse mixture‑of‑experts, and it can further improve accuracy through test‑time compute scaling.

By TabLDM Team, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang, Chunxiao Liu, Erli Meng, Bin Wang
arXiv Machine Learning
Aug 27

EXAONE Tabular 1.0 : Technical Report

EXAONE Tabular 1.0 is a compact tabular foundation model family that performs classification and regression via in-context learning without dataset-specific gradient updates. It is pretrained exclusively on a synthetic structural‑causal‑model prior and introduces an architecture‑centered redesign that interleaves feature‑axis and item‑axis attention within each Transformer layer, mediated by summary tokens. Across four public benchmarks, its 20.81 M‑parameter classification model ranks first on TabArena, surpassing tuned ensembles and AutoML pipelines, while its regression model matches the performance of a 1.64 B‑parameter model at roughly one‑eleventh the inference cost, and it achieves top rankings on BCCO, TALENT, and ScoringBench.

By Moonjung Eo, Min-Kook Suh, Hye-Seung Cho, Jiwon Kim, Seoyoon Kim, Sangjun Nam, Soonyoung Lee
arXiv AI
Jun 30

Beyond IID: How General Are Tabular Foundation Models, Really?

arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.

By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv AI
Jun 16

LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction

arXiv:2606. 15314v1 Announce Type: cross Abstract: Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take.

By Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou
arXiv Machine Learning
Aug 14

TabH2O: A Unified Foundation Model for Tabular Prediction

arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.

By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati