arXiv:2607. 11956v1 Announce Type: cross Abstract: Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is the version deployed in practice: the exact estimator shipped by toolkits such as pyDVL and OpenDataVal.
By Zongye Lyu
arXiv:2609.05877v1 Announce Type: new
Abstract: Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configu...
By Jia Bi, Alin-Marin Elena
arXiv:2606. 29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups.
By Louis Berthier, Ahmed Shokry, Maxime Moreaud, Guillaume Ramelet, Aymeric Dieuleveut
The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.
By Samir Orujov
arXiv:2607. 21003v1 Announce Type: new Abstract: Ordinal Classification (OC) deals with classification tasks where the classes follow a natural order.
By Rafael Ayll\'on-Gavil\'an, Francisco Jos\'e Mart\'inez-Estudillo, David Guijo-Rubio, C\'esar Herv\'as-Mart\'inez, Pedro A. Guti\'errez
arXiv:2605. 20716v5 Announce Type: replace Abstract: Random forests construct each tree with a different, randomised representation of the feature space.
By Youngjoon Park
arXiv:2606. 18853v1 Announce Type: cross Abstract: A recent line of work has reframed individual decision trees as linear models on engineered features associated with their splits, opening routes for oracle inequalities and feature-importance reinterpretation, but leaving open the question of what unified geometric object a forest induces when one indexes its feature map by nodes rather than by splits.
By Nicolas Mahler
arXiv:2606. 27997v1 Announce Type: new Abstract: Benchmarks of machine learning models often include many datasets, making evaluation expensive.
By Rostislav Gusev, Alexey Zaytsev
The paper introduces the Adaptive Margin Ordinal Loss (AMOL), a new loss function designed to reduce the tendency of neural networks to predict center classes in ordinal classification tasks—a problem called center‑class hedging. AMOL applies a multiplicative weight to per‑class loss terms that is large only when a candidate class is near the center while the true label is far from it, thereby discouraging hedging. The authors also propose the Center‑Hedging Rate (CHR) metric to quantify this failure mode and demonstrate that AMOL achieves state‑of‑the‑art Quadratic Weighted Kappa scores on four benchmarks, with an asymmetric variant eliminating hedging on the Abalone dataset.
By Manisha Kandel
arXiv:2606. 24903v2 Announce Type: replace Abstract: Few-shot label acquisition lacks a label-free signal for when additional labels cease to improve accuracy: existing stopping criteria either require a held-out validation set (violating the few-shot premise) or rely on theoretically ungrounded heuristics, so we introduce the spectral saturation index $S(K)=\mathrm{erank}(\hat{\Sigma}_W^{(K)})/K$, the exponential spectral entropy of the pooled within-class covariance normalized by per-class support size $K$, which measures the exploration rate per label and falls below a fixed threshold $\tau=0.
By Arnav Gupta
arXiv:2607. 26628v1 Announce Type: new Abstract: TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates.
By Mohammed Abdullah
arXiv:2606. 04161v1 Announce Type: new Abstract: Different predictors often excel on different inputs, so picking the best one per instance promises higher accuracy than committing to a single model.
By Tyler Crosse, Alan Nadelsticher Ruvalcaba, Dustin Khang LeDuc, Thomas Trask, Nicholas Lytle, David Joyner