arXiv Machine Learning

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

The paper introduces the Descriptive‑Complexity Information Criterion (DCIC), a new framework for selecting models when predictors are highly correlated and the model class is uncertain. DCIC uses Kraft‑admissible code lengths to regularize large collections of candidate models, achieving selection consistency under sub‑Weibull noise without requiring RIP‑type conditions and providing non‑asymptotic oracle risk bounds even when the model is misspecified. The approach also unifies heterogeneous model classes on a common complexity scale, enables class–model recovery under identifiability conditions, and offers a complexity‑guided search path that balances computational effort with statistical accuracy, as demonstrated by numerical experiments.

arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Jun 15

A Complexity Measure for Active Learning in Multi-group Mean Estimation

arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.

By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub
Hugging Face Trending Papers
Aug 13

Bagging Robustly Learns VC Classes with Linear Sample Complexity

We revisit the problem of learning predictors robust to adversarial examples at test-time. We prove that VC classes are adversarially robustly learnable with sample complexity linear in the VC dimension $d$, providing an exponential improvement over the previous upper bound of Montasser, Hanneke, and Srebro (2019).

arXiv Machine Learning
Jul 8

Boosting with List-Decodable Codes

arXiv:2607. 05791v1 Announce Type: cross Abstract: Boosting is a fundamental technique for generically improving the accuracy of learning algorithms (Schapire 1989).

By Addison Prairie, Li-Yang Tan
arXiv Machine Learning
Sep 4

Resolution-Aware Experimental Design under Partial Identifiability

The paper introduces Resolution-Aware Experimental Design (RAED), a method that selects experiments by minimizing the expected size of the nonempty structural candidate set while controlling false-exclusion rates. RAED is shown to preserve expected ordering under a composite Blackwell comparison and is implemented via a learned score-based approach with finite-sample nuisance-average and positive-tail calibration. Experiments on subsurface-flow, fluvial, and methane-oxidation benchmarks demonstrate RAED’s ability to resolve structural ambiguities and provide finite-sample guarantees for tail-sensitive nuisance risk.

By Sofianos Panagiotis Fotias