The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.
By Samir Orujov
arXiv:2608.21262v1 Announce Type: cross
Abstract: Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cut...
By Adam Noonan
arXiv:2607. 07717v1 Announce Type: new Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups.
By Ha-Hieu Pham, Hai-Dang Nguyen, Dang P. M. Cao, Thanh-Huy Nguyen, Min Xu, Trung-Nghia Le, Ulas Bagci, Huy-Hieu Pham
arXiv:2608. 07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized.
By Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li
arXiv:2606. 27385v1 Announce Type: new Abstract: The most widely used RANSAC variants score candidate models by counting inliers or summing per-point scores that saturate beyond a residual threshold.
By James Pritts, Felix Seegr\"aber, Kevin K\"oser
The paper introduces Difference‑Quotient Clustering (DQC) to address mean‑collapse in multimodal regression. DQC partitions data by minimizing intra‑cluster output‑vs‑input discrepancy, assigning each sample to the cluster with the lowest maximum contradiction ratio. The resulting cluster labels train a logits generator and conditional network, achieving lower minimum squared error on synthetic benchmarks compared to random labeling and mean‑collapse baselines.
By Huang Weiquan
arXiv:2609.36762v1 Announce Type: new
Abstract: Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This ass...
By Mitushi Goyal, Tarun S., Riddhanya Senapathi, Arun Raman
In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroups. We study this pre-deployment fairness problem as an audit question: after a long-tailed multi-label CXR model is converted from scores into decisions, who is missed?
arXiv:2607. 16811v4 Announce Type: replace Abstract: Drift detectors that work tend not to explain themselves, and drift detectors that explain themselves tend to fail in high dimension.
By Behnam Asadi
arXiv:1907.06994v2 Announce Type: replace-cross
Abstract: Mixtures of experts (MoE) are conditional mixture models in which both the mixing proportions and the component densities depend on the predi...
By Thin Nguyen-Van, Faicel Chamroukhi, Ha Hoang Van, Bao Tuyen Huynh
arXiv:2606. 30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood.
By Anish Acharya, Kris W Pan, Brian Verkhovsky
arXiv:2609. 30477v1 Announce Type: cross Abstract: Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential.
By Yordan P. Raykov, Max A. Little