arXiv Machine Learning By Manish Kumar, Anton Frederik Thielmann, Christoph Weisser, Benjamin S\"afken

From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning

Read the original on arXiv Machine Learning →

arXiv:2604. 05635v2 Announce Type: replace Abstract: Numerical preprocessing remains a critical component of tabular deep learning, as the representation of continuous features can strongly affect downstream performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Tabular Numeric Stretch Transformation

arXiv:2608. 09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scales, and statistical properties.

By Zihao Ye, Juyong Kim, Johnna Sundberg, Burak Varici, Pradeep Ravikumar
arXiv Machine Learning
Sep 2

Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data

The paper introduces In-Table Prediction (ITB), a self‑supervised task where deep neural networks learn to predict any column in a table from the remaining columns. It proposes a novel neural layer to handle missing continuous values, generates synthetic datasets with controlled column relationships, and evaluates three architectures—MLP, ResNet, and Transformer—showing that attention‑based Transformers perform best when ample training data and large embeddings are used. The study is limited to synthetic, small‑column tables and is presented as an initial investigation rather than a comprehensive real‑world analysis.

By Xiao Zhao, Daniela Oelke
arXiv Machine Learning
Sep 17

TabICLv2: A better, faster, scalable, and open tabular foundation model

TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.

By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan