The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.
By Maximilian Schambach, Clemens Biehl, Sam Thelin
Support-Compiled Feature Folding (SCFF) is a training‑free inference framework that addresses the feature‑side scaling dilemma in tabular foundation models by routing support‑ranked features through bounded leaves of the native encoder, checking residual evidence, and merging encoded messages for a single contextual prediction. This approach transforms quadratic pairwise mixing into linear‑in‑width work with a bounded local working set, achieving dataset‑macro accuracy and NLL improvements across six backbones on an 18‑dataset wide‑table slice. SCFF delivers significant GPU‑memory savings (median 2.09×–2.36×) and, when constrained by a peak‑memory ceiling, further boosts accuracy by up to 4.06 points over the widest single‑leaf baseline.
whyItMatters":"SCFF demonstrates that memory‑efficient inference can simultaneously improve accuracy and reduce resource usage in tabular foundation models, offering a practical solution for deploying these models at scale."
By Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
The paper introduces RefineICL, an attention‑gated, feed‑forward‑network‑free framework that refines representations in situ for tabular foundation models. By using support labels to guide episode‑specific updates, the method transfers learned corrections to unlabeled queries without altering model parameters, achieving state‑of‑the‑art performance on AMLB29 and TabArena benchmarks. Experiments and internal interventions demonstrate that intermediate support updates are essential for constructing task‑specific predictors in context.
By Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
arXiv:2605.12904v2 Announce Type: replace
Abstract: Tabular foundation models (TFMs) have emerged as a powerful paradigm for in-context learning on structured data, enabling direct prediction on new...
By Yilong Chen, Xueying Ding, Leman Akoglu
TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.
By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
arXiv:2607. 17419v1 Announce Type: cross Abstract: Linear attention promises constant-time recurrent inference but degrades sharply on associative recall.
By Ayoub Ghriss, Sourav Chakraborty
Causilo is a new tabular foundation model that delivers state‑of‑the‑art predictive performance while achieving exceptionally fast inference. On the TabArena benchmark it scores 1785.4 Elo with a median inference time of 0.10 seconds per 1 K test samples, outperforming TabPFN‑3.5‑Fast by 31.6% in speed and reaching the performance–efficiency Pareto frontier. The architecture builds on TabICL’s column‑then‑row design, adding a row‑refinement module that exchanges information among cell representations before a final column stage, and uses cross‑attention with a fixed number of summary tokens to keep attention cost linear in the number of features.
By Minyong Cho, Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv:2606. 07345v1 Announce Type: new Abstract: Tabular foundation models, exemplified by TabPFN, perform prediction via in-context learning, inferring test labels directly from labeled training examples.
By Si-Yang Liu, Han-Jia Ye
arXiv:2609.36883v1 Announce Type: new
Abstract: Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without...
By Tianqi Zhao, Tianyi Zhuang, Shuo Duan, Guanyang Wang, Yan Shuo Tan, Qiong Zhang
CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling) is a new method for long-context LLM inference that replaces costly quadratic attention prefilling with a dynamic, input-adaptive sparse routing scheme. It introduces a structural proxy, C_struct, to directly read routing decisions from the proxy attention map, eliminating the need for pooled matrix multiplication and KL divergence. Additionally, CRISP addresses the post-softmax mass cliff by using a sink-aware threshold based on the noise floor, theoretically reducing background noise accumulation to O(n). Empirical results on InfiniteBench, RULER, and LongBench show that CRISP outperforms existing sparse methods and can match or exceed exact dense attention, achieving up to a 5.30× speedup at 512k tokens and significant gains on retrieval-heavy tasks.
By Huu Huy Nguyen, Chien Van Nguyen, Franck Dernoncourt, Ryan A. Rossi, Linh Ngo Van, Jieyang Chen, Thien Huu Nguyen
arXiv:2607. 28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs.
By Haozhe Hu, Hao Wu, Peiran Yin, Chao Han, Yunpu Ma, Xiaoyu Shen
arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.
By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao