arXiv Machine Learning By Yanhang Zhang, Wei Liu, Yuhong Yang

A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs

Read the original on arXiv Machine Learning →

The paper introduces the Descriptive‑Complexity Information Criterion (DCIC), a new framework for selecting models when predictors are highly correlated and the model class is uncertain. DCIC uses Kraft‑admissible code lengths to regularize large collections of candidate models, achieving selection consistency under sub‑Weibull noise without requiring RIP‑type conditions and providing non‑asymptotic oracle risk bounds even when the model is misspecified. The approach also unifies heterogeneous model classes on a common complexity scale, enables class–model recovery under identifiability conditions, and offers a complexity‑guided search path that balances computational effort with statistical accuracy, as demonstrated by numerical experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Jun 15

A Complexity Measure for Active Learning in Multi-group Mean Estimation

arXiv:2606. 14690v1 Announce Type: new Abstract: We study a \emph{max-risk} objective for active learning in a multi-group mean estimation $d$-armed bandits: a learner adaptively allocates a budget of $T$ samples across $d$ groups to minimize the worst-case uncertainty index $\max_{k\in[d]}\sigma_k^2/n_k$, where $\sigma_k$ is the standard deviation of the distribution of arm $d$, and $n_k$ is the number of times arm $d$ is sampled.

By Abdellah Aznag, Rachel Cummings, Adam N. Elmachtoub