arXiv:2512. 22240v5 Announce Type: replace-cross Abstract: Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biological findings.
By Chama Bensmail
arXiv:2506. 01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.
By Jelke Wibbeke, Sebastian Rohjans, Andreas Rauh
arXiv:2512. 13003v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is essential for determining when a supervised model encounters inputs that differ meaningfully from its training distribution.
By Min Lu, Hemant Ishwaran
arXiv:2512. 08216v4 Announce Type: replace-cross Abstract: Accurate segmentation of lung tumors from 3D computed tomography (CT) scans is essential for automated treatment planning and response assessment.
By Aneesh Rangnekar, Harini Veeraraghavan
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen
arXiv:2603. 24025v2 Announce Type: replace Abstract: Unsupervised learning of high-dimensional data is challenging due to irrelevant or noisy features obscuring underlying structures.
By Chen Ma, Wanjie Wang, Shuhao Fan
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
arXiv:2606. 00327v1 Announce Type: cross Abstract: Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries.
By Kai R. Wycik, Tiffany M. Tang, Tarek M. Zikry, Genevera I. Allen
arXiv:2506. 22427v2 Announce Type: replace-cross Abstract: We propose CLoVE (Clustering of Loss Vector Embeddings), a novel algorithm for Clustered Federated Learning (CFL).
By Randeep Bhatia, Nikos Papadis, Murali Kodialam, TV Lakshman, Sayak Chakrabarty
Interpretable AI with Local Distillation proposes a method where a black‑box teacher model guides a regularized linear student model at each query point. The teacher defines locality by upweighting training observations with similar predicted outcomes and anchors the fit with its own prediction at the query point, treated as a pseudo‑observation. By adding Gaussian randomization and refitting, the approach identifies reliable features and stable subgroups, achieving near‑teacher accuracy while producing sparse, locally interpretable linear models.
By Erin Craig, Yiling Huang, Snigdha Panigrahi
CORE-STACK+ is a new meta‑learning framework for deep stacked generalization that tackles two key problems in heterogeneous vision ensembles: prediction‑space multicollinearity and calibration collapse. It introduces a four‑step preconditioning pipeline—kernelized redundancy filtering, a lightweight differentiable meta‑feature gate, a spectrum‑adaptive ridge penalty, and a Laplace‑approximate Bayesian blender—to jointly improve conditioning and calibration. Across six vision benchmarks, CORE‑STACK+ boosts accuracy, reduces model count and inference cost, and significantly lowers expected calibration error compared to existing methods.
By Noor Islam S. Mohammad
arXiv:2607. 16250v1 Announce Type: cross Abstract: Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection.
By Priyanka Paudel, Madan Baduwal