arXiv:2609.31019v2 Announce Type: replace-cross
Abstract: Trimming methods for robust clusterwise regression discard a fixed fraction of the data. Too low a level breaks the fit; too generous a level...
By Samir Orujov
arXiv:2608.21262v1 Announce Type: cross
Abstract: Many machine-learning systems set a threshold at a quantile of a calibration set: conformal predictors that promise 90% coverage by drawing their cut...
By Adam Noonan
arXiv:2609.36762v1 Announce Type: new
Abstract: Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This ass...
By Mitushi Goyal, Tarun S., Riddhanya Senapathi, Arun Raman
arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2606. 16341v1 Announce Type: new Abstract: A filtered approximate-nearest-neighbor (ANN) query returns the k nearest vectors among those satisfying an attribute predicate P of selectivity s.
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2511.22435v2 Announce Type: replace
Abstract: Invariant learning on graphs aims to build predictors that rely on causal substructures rather than on environment-specific shortcuts. Current meth...
By Ali Ghasemi, Farooq Ahmad Wani, Maria Sofia Bucarelli, Fabrizio Silvestri
arXiv:2606. 29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups.
By Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric Dieuleveut
arXiv:2504. 07133v2 Announce Type: replace-cross Abstract: We revisit the problem of estimating $k$ linear regressors with self-selection bias in $d$ dimensions with the maximum selection criterion, as introduced by Cherapanamjeri, Daskalakis, Ilyas, and Zampetakis [CDIZ23, STOC'23].
By Alkis Kalavasis, Anay Mehrotra, Felix Zhou
arXiv:2607. 01945v1 Announce Type: cross Abstract: The classical $k$-means clustering cannot be directly used to incomplete data, and existing $k$-means-based clustering for missing data primarily focus on improving the practical accuracy of clustering, whereas most of them lack theoretical guarantees in the asymptotic sense.
By Xin Guan
The paper introduces Difference‑Quotient Clustering (DQC) to address mean‑collapse in multimodal regression. DQC partitions data by minimizing intra‑cluster output‑vs‑input discrepancy, assigning each sample to the cluster with the lowest maximum contradiction ratio. The resulting cluster labels train a logits generator and conditional network, achieving lower minimum squared error on synthetic benchmarks compared to random labeling and mean‑collapse baselines.
By Huang Weiquan
arXiv:2606. 06469v1 Announce Type: cross Abstract: Let $S$ be the set of unit norm linear classifiers $\theta \in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $\kappa$ fixed in advance.
By August Y. Chen, Ahmed El Alaoui
arXiv:2608. 08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may have to answer queries whose coordinate and inspection time are chosen only after the data are seen.
By Ibne Farabi Shihab, Adria Binte Habib