arXiv Machine Learning

Impact of canny edge detection preprocessing on performance of machine learning models for Parkinson's disease classification

arXiv Machine Learning
4d ago

Vectorized Dynamic Histograms for Sparse Oblique Forests

The paper presents optimizations for Sparse Oblique (SPO) forests in Google’s Yggdrasil Decision Forests, addressing training speed issues caused by runtime sampling of sparse linear feature combinations. By fixing inefficiencies and introducing two new methods—hierarchical AVX2/AVX-512 vectorized histogram filling and runtime‑dynamic histograms—the authors achieve 2–5× speedups for both Gradient Boosted Trees and Random Forests, bringing SPO‑RF training time on par with axis‑aligned RFs. Extensive evaluation on 19 datasets, including up to 10.5 million rows and 1.6 million features, demonstrates these improvements without compromising accuracy.

By Ariel Lubonja, Jungsang Yoon, Haoyin Xu, Yue Wan, Yilin Xu, Richard Stotz, Mathieu Guillame-Bert, Joshua T. Vogelstein, Randal Burns
arXiv Machine Learning
Jun 16

Machine Learning and the Random Walk Puzzle: Forecasting the CAD/USD Exchange Rate with Expanding Window Evaluation and SHAP Interpretability

arXiv:2606. 15058v1 Announce Type: new Abstract: This study examines whether machine learning (ML) models can outperform the naive random walk benchmark in forecasting the monthly USD/CAD exchange rate.

By Louis Agyekum, Edmund Fosu Agyemang, Obu-Amoah Ampomah, Kofi Acheampong, Emmanuel Boadi, Priscilla Yaa Amakye, Fafa Shalom Tchorly, Enock Adu Bonsu, Eric Nyarko
arXiv AI
Jun 16

LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction

arXiv:2606. 15314v1 Announce Type: cross Abstract: Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take.

By Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou
arXiv Machine Learning
Aug 14

TabH2O: A Unified Foundation Model for Tabular Prediction

arXiv:2605. 18383v2 Announce Type: replace Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning.

By Pascal Pfeiffer, Dmitry Gordeev, Mathias M\"uller, Laura Fink, Joan Salv\`a Soler, Mark Landry, Branden Murray, Marcos V. Conde, Sri Satish Ambati