arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.
By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz
arXiv:2607. 16250v1 Announce Type: cross Abstract: Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection.
By Priyanka Paudel, Madan Baduwal
The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.
By Bilal Ahmad, Rajed Mehmood
arXiv:2609.09189v1 Announce Type: cross
Abstract: High classification accuracy alone is insufficient for clinical image analysis, where calibrated confidence and reliable uncertainty estimates are es...
By Nisreen Albzour, Sarah S. Lam
arXiv:2607. 17345v1 Announce Type: new Abstract: Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options.
By Mohammed Saeed Al-Huraibi, Ihsan Yozgat, Ahmet Kaplan
arXiv:2601. 05151v3 Announce Type: replace-cross Abstract: Feature selection (FS) is essential for biomarker discovery and clinical predictive modeling.
By Anastasiia Bakhmach, Paul Dufoss\'e, Simon Charpigny, Florence Monville, Laurent Greillier, Fabrice Barl\'esi, S\'ebastien Benzekry
The paper benchmarks six long‑tail loss functions—cross‑entropy, weighted CE, class‑balanced loss, focal loss, LDAM, and logit‑adjusted softmax—across three single‑cell foundation model architectures (scGPT, scBERT, Geneformer) and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas). It shows that overall accuracy masks systematic failures on rare, disease‑relevant cell types, with a consistent gap between overall accuracy, Macro‑F1, and rare‑class recall under plain cross‑entropy. The study identifies two distinct regimes of rare‑class failure, predicts reweighting efficacy by absolute training‑set size, and finds class‑balanced loss and LDAM to be the most reliable across all settings.
By Zeyu Dong, Jiahui Zhong
arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.
By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv:2512. 17678v2 Announce Type: replace-cross Abstract: Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling.
By Daphn\'e Chopard, Jorge da Silva Gon\c{c}alves, Irene Cannistraci, Thomas M. Sutter, Julia E. Vogt
The study evaluates how different molecular feature spaces—Morgan fingerprints, RDKit physicochemical descriptors, and SMILES bigrams—affect the prediction of blood‑brain barrier permeability using various learning algorithms. Dynamic Random Forests with combined features achieved the best performance (mean AUC 0.970). When applying Generalized Random Forests to estimate heterogeneous effects of LogP on BBB permeability, orthogonalization revealed that apparent heterogeneity largely vanished after accounting for confounding, shifting importance toward residual SMILES bigram information.
By Tshemollo Rapolai, Seite Makgai, Mohammad Arashi
EMFE (Efficient Mathematical Feature Extraction) is a lightweight, explainable machine‑learning framework that classifies single red‑blood‑cell images as parasitized or uninfected using five engineered features: Gray World color normalization, adaptive green‑channel thresholding, morphological spot detection, and classical classifiers. On the NIH LHNCBC malaria dataset (27,558 images from 200 patients), a tuned Random Forest achieved 94.6% pooled out‑of‑fold accuracy, 94.3% on a 40‑patient holdout, and outperformed deep‑learning baselines in an accuracy‑efficiency trade‑off. Ablation studies, synthetic perturbations, and explainability analyses identified spot saturation as the dominant discriminative feature and quantified the framework’s failure modes and patient‑level performance.
By Md Abdullah Al Kafi, Walayat Hussain, Mousumi Karmakar, Sumit Kumar Banshal, Ahmed Al Marouf