arXiv AI

When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

arXiv:2606. 19827v1 Announce Type: cross Abstract: Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often require costly expert adjudication, even though structured clinical variables are routinely available in tabular form.

arXiv AI
Sep 10

TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables

TabBench-Bio is a living, interactive benchmark that evaluates machine learning models on 43 high‑dimensional biomedical tables, covering multiple domains. Using a shared cross‑validation protocol, the benchmark compares classical estimators, neural networks, and tabular foundation models across 28 feature‑by‑sample operating points, with RealTabPFN v2.5 achieving the highest performance at the reference cell of 10,000 features and 100 training samples. The benchmark provides reproducible results, fold‑level predictions, and invites community contributions to expand its dataset collection.

By Jules Kreuer, Sofiane Ouaari, Julia Hellmig, Julius Braitinger, Nico Pfeifer
arXiv AI
Aug 11

Tabular Numeric Stretch Transformation

arXiv:2608. 09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties.

By Zihao Ye, Juyong Kim, Johnna Sundberg, Burak Varici, Pradeep Ravikumar