arXiv Statistics ML

Minimax Additive Regression under Unknown Dependent Designs

arXiv:2609. 39212v1 Announce Type: new Abstract: We study additive regression under a potentially non-product random design on $[0,1]^d$, allowing the dimension $d$ to grow with the sample size $n$.

arXiv Machine Learning
Aug 27

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

The paper introduces a pooling‑ridge estimation method for functional linear regression that handles data observed at discrete times, ranging from sparse to dense designs. By combining pooling strategies with RKHS‑based techniques, the authors achieve minimax‑optimal prediction risk for both scalar‑on‑function and function‑on‑function models. The study identifies distinct phase transitions in convergence behavior, with up to three transitions for function‑on‑function regression, and validates the approach through simulations and real data examples.

By Shunxing Yan, Fang Yao
arXiv Statistics ML
Sep 7

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The paper investigates Double Machine Learning (DML) estimators under structure‑agnostic (SA) models, which assume the data‑generating law lies within a neighborhood of fixed machine‑learning estimates. It shows that for two of three studied functionals—the quadratic functional in the Gaussian sequence model and the quadratic density integral functional—the DML estimators are asymptotically inadmissible, being dominated by second‑order empirical higher‑order influence function (HOIF) estimators. For the third functional, the expected conditional covariance, both DML and HOIF estimators remain minimax but neither dominates the other.

By Lin Liu, Rajarshi Mukherjee, James M Robins
arXiv Machine Learning
Aug 27

Generalized Riesz Regression: A Unified Framework for Debiased Machine Learning with Riesz Representer Fitting under Bregman Divergence

The paper introduces generalized Riesz regression, a framework that minimizes a Bregman divergence made observable through the Riesz identity. By selecting squared or Kullback–Leibler-type divergences, it recovers existing Riesz regression, tailored loss minimization, and density‑ratio objectives. The authors derive first‑order conditions that enforce empirical Riesz equations in model‑dependent tangent directions, provide convergence rates for sparse, RKHS, and neural network models, and establish asymptotic normality under Donsker or cross‑fitting conditions, with applications to treatment effects, average marginal effects, and covariate shift.

By Masahiro Kato
arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Machine Learning
Jun 4

Orthogonal Learner for Estimating Heterogeneous Long-Term Treatment Effects

arXiv:2604. 00915v2 Announce Type: replace Abstract: Estimation of heterogeneous long-term treatment effects (HLTEs) is relevant for personalized decision-making in marketing, economics, and medicine, where short-term observational datasets are often combined with long-term observational datasets.

By Haorui Ma, Dennis Frauen, Valentyn Melnychuk, Stefan Feuerriegel
arXiv Statistics ML
Sep 11

Learning-Based Surrogate Method for Stochastic Optimization under Decision-Dependent Uncertainty with Adaptive Random Designs

The paper introduces a learning-based surrogate approach for stochastic optimization problems where uncertainty depends on the decision, modeled via a nonparametric regression. It constructs a surrogate that embeds iteratively updated Jacobian estimates, using an adaptive random design that focuses sampling near the current iterate to achieve dimension‑independent convergence of the Jacobian estimates. The resulting learning‑based stochastic prox‑linear (L‑SPL) algorithm demonstrates nonasymptotic convergence rates and outperforms existing methods in sample efficiency and objective value in numerical experiments.

By Boyang Shen, Junyi Liu