arXiv Machine Learning By Maximilian Schambach, Clemens Biehl, Sam Thelin

Benchmarking Attention for Tabular Foundation Models

Read the original on arXiv Machine Learning →

The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Block Sparse Flash Attention

Block Sparse Flash Attention (BSFA) is a drop‑in replacement for FlashAttention that speeds up long‑context inference by pruning about 50% of computation and memory transfers. It selects the top‑k most important value blocks for each query using exact query‑key similarities and calibrated per‑layer, per‑head thresholds, requiring only a one‑time training‑free calibration. On Llama‑3.1‑8B, BSFA delivers up to 1.13× speedup on LongBench with a 1.1% accuracy drop and up to 1.24× on Needle‑in‑a‑Haystack retrieval with a 1% drop, while the attention kernel itself accelerates by up to 1.38×.

By Daniel Ohayon, Itay Lamprecht, Itay Hubara, Israel Cohen, Daniel Soudry, Noam Elata
arXiv Machine Learning
Sep 22

Causilo Technical Report

Causilo is a new tabular foundation model that delivers state‑of‑the‑art predictive performance while achieving exceptionally fast inference. On the TabArena benchmark it scores 1785.4 Elo with a median inference time of 0.10 seconds per 1 K test samples, outperforming TabPFN‑3.5‑Fast by 31.6% in speed and reaching the performance–efficiency Pareto frontier. The architecture builds on TabICL’s column‑then‑row design, adding a row‑refinement module that exchanges information among cell representations before a final column stage, and uses cross‑attention with a fixed number of summary tokens to keep attention cost linear in the number of features.

By Minyong Cho, Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
arXiv Machine Learning
Aug 31

SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning

SOMTab is a Set-Order Mamba architecture designed for efficient tabular in-context learning. It separates representation construction from query-conditioned retrieval, using Mamba-based state‑space mixing to build compact row and column representations while retaining attention for final prediction. The model, along with a synthetic prior called DCH‑TailMix, achieves performance comparable to Transformer‑based tabular foundation models but with faster inference and lower GPU memory usage.

By Hao Wang, Siyu Zhang, Wei Ma
arXiv Machine Learning
Sep 14

Attention Quantization for Tabular Foundation Models

The paper introduces a quantization technique for tabular foundation models that focuses on converting queries, keys, and values to FP8 and employing explicit FP8 matrix multiplication to accelerate attention calculations. It emphasizes aligning quantization errors between training and test rows to avoid accuracy loss, and demonstrates up to 1.7× speedup over 16‑bit kernels with no significant accuracy degradation on TabPFN‑v3 and TabICLv2 across TabArena and BeyondArena.

By Jonas M. K\"ubler, Benjamin J\"ager, Klemens Fl\"oge, Noah Hollmann, Frank Hutter