arXiv AI

Bounded Context Management for Tabular Foundation Models on Stream Learning

arXiv:2606. 18677v1 Announce Type: cross Abstract: Tabular stream learning requires predictions on sequentially arriving examples under distribution shift.

arXiv Machine Learning
Aug 31

PolicyLong: Towards On-Policy Context Extension

PolicyLong introduces a dynamic on‑policy approach to constructing long‑context data for large language models, addressing the off‑policy gap of previous single‑pass methods. By repeatedly re‑screening data using the model’s current entropy landscape, it creates a self‑curriculum that aligns training distribution with evolving model capabilities. Experiments on RULER, HELMET, and LongBench‑v2 demonstrate consistent performance gains, especially at longer contexts, outperforming EntropyLong and NExtLong.

By Junlong Jia, Jiang Zhou, Ziyang Chen, Xing Wu, Chaochen Gao, TingHao Yu, Feng Zhang, Songlin Hu
arXiv Machine Learning
Aug 24

Tydra: An Efficient Hybrid Model for Tabular Data

Tydra is a hybrid Transformer‑State Space Model that interleaves attention and SSM layers for tabular in‑context learning. It achieves a 30% reduction in inference time compared to the Transformer‑only TabPFN while preserving most of its predictive performance. On 30 OpenML datasets, Tydra also outperforms a Hydra model that is roughly ten times larger, demonstrating that hybrid architectures can balance accuracy and efficiency for tabular foundation models.

By Mieszko Komisarczyk, Saurabh Mathur, Maurice Kraus, Sriraam Natarajan, Kristian Kersting
arXiv Machine Learning
4d ago

TabFM: A Zero-Shot Foundation Model for Tabular Data

arXiv:2609.37959v1 Announce Type: new Abstract: Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We p...

By Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
arXiv Machine Learning
Aug 31

SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning

SOMTab is a Set-Order Mamba architecture designed for efficient tabular in-context learning. It separates representation construction from query-conditioned retrieval, using Mamba-based state‑space mixing to build compact row and column representations while retaining attention for final prediction. The model, along with a synthetic prior called DCH‑TailMix, achieves performance comparable to Transformer‑based tabular foundation models but with faster inference and lower GPU memory usage.

By Hao Wang, Siyu Zhang, Wei Ma