arXiv AI

Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research

arXiv Machine Learning
Aug 18

On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

arXiv:2608. 14705v1 Announce Type: cross Abstract: Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging.

By Ljubomir Buturovic (East Palo Alto, United States)
arXiv AI
Aug 14

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

arXiv:2608. 12446v1 Announce Type: cross Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability.

By Seyyed Ali Hoseini, Javad Baseri, Hamid Saadatfar, Edris Hoseini Gol, AmirHossein Eshghi
arXiv AI
3d ago

NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

NeuroAtlas is the largest EEG benchmark to date, comprising 42 datasets and 260,000 hours of clinical EEG data across epilepsy, sleep medicine, brain age estimation, and brain‑computer interfaces. The study evaluates foundation models (FMs) for EEG against supervised baselines and generic time‑series FMs, finding that EEG‑specific FMs do not consistently outperform generic ones. It also demonstrates that standard machine‑learning metrics are inadequate for clinical relevance, advocating for task‑specific measures such as event‑level decision quality, hypnogram features, and brain‑age gap.

By Konstantinos Kontras, Trui Osselaer, Stylianos G. Mouslech, Angeliki-Ilektra Karaiskou, Guido Gagliardi, Thomas Strypsteen, Mohammad Hossein Badiei, Anku Rani, Maarten Vanmarcke, Miguel Bhagubai, Chanakya Ekbote, Jaedong Hwang, Christos Chatzichristos, Paul Pu Liang, Maarten De Vos
arXiv Machine Learning
Sep 4

RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models

RobustSeiz is an open‑source, model‑agnostic framework designed to benchmark the robustness of EEG seizure detection models under realistic clinical stressors. It standardizes four public scalp‑EEG corpora into BIDS‑EEG trees, applies controlled distribution shifts—including environmental, noise, and adversarial transforms—across predefined hyperparameter grids, and reports comprehensive performance metrics such as sensitivity, precision, F1, false positives per 24 h, onset timing, and predictive agreement. The framework offers a Dockerized GPU pipeline, experiment registry, and both full‑evaluation and research‑subset modes, and demonstrates its utility by evaluating a contemporary detector on TUSZ across the full shift grid.

By Mohammad Mohammadi, Alireza Zarei
arXiv Machine Learning
Sep 25

Limited Structural Reliability in Public Educational Prediction Benchmarks: A Four-Dimension Audit of Seven Datasets

The study audited seven public educational prediction datasets using four pre‑modeling reliability checks—baseline gap, split instability, null separation, and metadata adequacy under group‑aware holdout. Only three datasets passed all checks; the others failed either group‑aware generalization tests or lacked necessary provenance metadata. The audit revealed that cross‑group fragility, rather than weak iid performance, was the dominant failure mode, and that increasing model complexity did not resolve these structural issues.

By Yan Ma, Lizhuo Zhang