The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.
By Maximilian Schambach, Clemens Biehl, Sam Thelin
The paper introduces TEmBed, a unified benchmark for evaluating tabular embeddings across four representation levels—cell, row, column, and table—using a diverse set of models. It demonstrates that the best model depends on the specific task and representation level, providing practical guidance for selecting embeddings in real-world applications. The study aims to facilitate the development of more general-purpose tabular representation models.
By Liane Vogel, Kavitha Srinivas, Niharika D'Souza, Sola Shirai, Oktie Hassanzadeh, Horst Samulowitz
arXiv:2607. 24130v1 Announce Type: cross Abstract: Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction.
By Ayeen Poostforoushan, Liane Vogel, Carsten Binnig
arXiv:2606. 30336v1 Announce Type: new Abstract: We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders.
By Marek Polewczyk, Maximilian Schambach, Marco Spinaci, Sam Thelin, Johannes H\"ohne
TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.
By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
LoopICL is a transformer architecture that loops a single block to address tabular tasks. It separates parameter count from computational depth by using a cell stream for per‑cell features and a row stream for in‑context examples, refined via within‑column and cross‑column attention. During pre‑training, varying loop counts and a learned exit‑gate allow the model to adjust inference depth at test time, achieving competitive performance with TabICLv2 while using about 90% fewer parameters.
By Amir Rezaei Balef, Katharina Eggensperger