arXiv Statistics ML

Ensembles of Exactly Solved Subsamples for Clusterwise Regression: Trimming Without a Trimming Level

The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.

arXiv Machine Learning
Aug 27

Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment

The paper introduces Difference‑Quotient Clustering (DQC) to address mean‑collapse in multimodal regression. DQC partitions data by minimizing intra‑cluster output‑vs‑input discrepancy, assigning each sample to the cluster with the lowest maximum contradiction ratio. The resulting cluster labels train a logits generator and conditional network, achieving lower minimum squared error on synthetic benchmarks compared to random labeling and mean‑collapse baselines.

By Huang Weiquan
arXiv Machine Learning
Jun 5

How abundant are good interpolators?

arXiv:2606. 06469v1 Announce Type: cross Abstract: Let $S$ be the set of unit norm linear classifiers $\theta \in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $\kappa$ fixed in advance.

By August Y. Chen, Ahmed El Alaoui