arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
arXiv:2605. 25050v2 Announce Type: replace-cross Abstract: Integrating multimodal datasets in clinical oncology is frequently hindered by high dimensionality and blockwise missingness, where entire data sources are unavailable for specific patient subsets.
By Mohamed Boussena, Florence Monville, Jacques Fieschi-Meric, Frederic Vely, Pierre Milpied, Julien Mazieres, Maurice Perol, Eric Vivier, Laurent Greillier, Fabrice Barlesi, Sebastien Benzekry
The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.
By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz
arXiv:2603.02221v3 Announce Type: replace-cross
Abstract: In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasi...
By Zizheng Zhang, Yiming Li, Justin Xu, Jinyu Wang, Rui Wang, Lei Song, Jiang Bian, David W Eyre, Jingjing Fu
arXiv:2607. 17345v1 Announce Type: new Abstract: Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options.
By Mohammed Saeed Al-Huraibi, Ihsan Yozgat, Ahmet Kaplan
arXiv:2607. 01001v1 Announce Type: cross Abstract: Radiomics is the established approach for CT-based lung cancer phenotyping, yet comparisons with foundation models rarely isolate contributions of feature extractor, classification head, and segmentation choice, or test cross-cohort robustness.
By Nils Neukirch, Martin Maurer, Nils Strodthoff