arXiv:2607. 11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal.
By Zongye Lyu
arXiv:2609.05877v1 Announce Type: new
Abstract: Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configu...
By Jia Bi, Alin-Marin Elena
arXiv:2606. 29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups.
By Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric Dieuleveut
The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.
By Samir Orujov
arXiv:2607. 21003v1 Announce Type: new Abstract: Ordinal Classification (OC) deals with classification tasks where the classes follow a natural order.
By Rafael Ayll\'on-Gavil\'an, Francisco Jos\'e Mart\'inez-Estudillo, David Guijo-Rubio, C\'esar Herv\'as-Mart\'inez, Pedro A. Guti\'errez
arXiv:2605. 20716v5 Announce Type: replace Abstract: Random forests construct each tree with a different, randomised representation of the feature space.
By Youngjoon Park