arXiv Machine Learning By Jannis Maier, Lennart Purucker

HAPEns: Hardware-Aware Post-Hoc Ensembling for Tabular Data

Read the original on arXiv Machine Learning →

arXiv:2603. 10582v2 Announce Type: replace Abstract: Ensembling is commonly used in machine learning on tabular data to boost predictive performance and robustness, but larger ensembles often lead to increased hardware demand.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 17

TabICLv2: A better, faster, scalable, and open tabular foundation model

TabICLv2 is a new state‑of‑the‑art tabular foundation model that outperforms existing methods on regression and classification tasks. It relies on a synthetic data generation engine for diverse pretraining, architectural innovations such as a scalable softmax attention, and optimized training protocols that replace AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 surpasses the current best model, RealTabPFN‑2.5, without any tuning, while also being faster and capable of handling million‑scale datasets with limited GPU memory.

By Jingang Qu, David Holzm\"uller, Ga\"el Varoquaux, Marine Le Morvan
arXiv Machine Learning
5d ago

Tabular Imbalanced Learning: A Survey, Benchmark, and Practical Guide

The paper surveys tabular imbalanced learning and introduces TILBench, a benchmark evaluating over 40 methods on 57 datasets. It presents a unified taxonomy of approaches and shows that no single method dominates across all settings, with performance depending on dataset regimes and computational constraints. Practical recommendations for method selection and future research directions are provided.

By Ruizhe Liu, Jiaqi Luo