arXiv Machine Learning

When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares

The paper extends prior work on volume sampling by providing a Loewner envelope for the centered coefficient covariance in least‑squares regression with a fixed pool of features and responses. It characterizes when this envelope is tight, linking tightness to strict spectral properties of residuals, and introduces a residual‑augmented change of measure to derive a one‑sided slack bound. The results also offer geometric insights at the boundary and demonstrate non‑vacuous certificates through frozen‑feature examples, focusing on conditional centered, full‑Gram‑whitened covariance rather than population generalization.

arXiv Machine Learning
Sep 10

A Geometric Phase Boundary for Volume-Sampled Linear Readouts

The paper investigates volume‑sampled linear readouts with fixed feature pools and responses, focusing on the randomness introduced solely by subset selection. It establishes a globally sharp Loewner envelope for centered, full‑Gram‑whitened coefficient covariance and characterizes when a positive geometric margin exists versus when it vanishes, using conditions on residual covariance slack and a pairwise Naimark‑complement minor test. The results provide explicit geometric boundaries and conservative certificates for strictness and variance terms in fixed‑query squared loss, offering a design‑specific phase characterization for this randomized linear‑readout primitive.

By Kihun Rhee
arXiv Machine Learning
Sep 4

Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails

The paper investigates restricted eigenvalue (RE) bounds for norm‑regularized estimators under heavy‑tailed designs. It shows that the previously conjectured sample‑size law based on Gaussian width fails for heavy‑tailed measurements, due to a phenomenon called simultaneous threshold occupancy. The authors provide explicit counterexamples, derive worst‑case sample‑complexity bounds, and compare the behavior of heavy‑tailed versus Gaussian designs on constant‑width polyhedral descent cones.

By Shi Fu, Huibo Xu, Qixin Zhang, Dacheng Tao
arXiv Machine Learning
Jun 9

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

arXiv:2606. 08517v1 Announce Type: new Abstract: Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously upper-bounds the selected risk, lower-bounds the acceptance probability $\pacc$ above a floor $\pmin$, and lower-bounds the deployment utility.

By Xiaoli Yu, Jiamiao Liu
arXiv Machine Learning
Aug 20

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

The paper introduces “ℝD_{CF5}”, a probe‑based estimator that predicts the region‑wise gain of a dynamic ensemble over the best static blend in regression tasks under distribution shift. Across 12 benchmark dataset‑shift pairs, the estimator achieves a Spearman correlation of +0.98 with actual test gains, outperforming alternative diagnostics. The authors also present a Probe‑Validated Ensemble Selector that chooses between a static affine stacker and dynamic realizers, demonstrating risk reductions of up to 16% in prospective deployments.

By Tianxin Zhou, Ruixi Lin
arXiv Machine Learning
3d ago

Structured Features Overfit Where Random Features Grok

arXiv:2609. 15047v1 Announce Type: new Abstract: Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/\lambda$ in the weight decay.

By Chon-Fai Kam, Miloud Bessafi, Frederic Cadet