arXiv Machine Learning By Masahiro Kato

Generalized Riesz Regression: A Unified Framework for Debiased Machine Learning with Riesz Representer Fitting under Bregman Divergence

Read the original on arXiv Machine Learning →

The paper introduces generalized Riesz regression, a framework that minimizes a Bregman divergence made observable through the Riesz identity. By selecting squared or Kullback–Leibler-type divergences, it recovers existing Riesz regression, tailored loss minimization, and density‑ratio objectives. The authors derive first‑order conditions that enforce empirical Riesz equations in model‑dependent tangent directions, provide convergence rates for sparse, RKHS, and neural network models, and establish asymptotic normality under Donsker or cross‑fitting conditions, with applications to treatment effects, average marginal effects, and covariate shift.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations

arXiv:2609. 03129v1 Announce Type: cross Abstract: Several classical machine-learning methods, such as KRRs and SVRs, are both computationally and analytically tractable since their estimators either admit closed-form expressions or are obtained by minimizing convex training objectives; neither feature is generally available for deep neural networks.

By Ruiyang Hong, Hrad Ghoukasian, Anastasis Kratsios
arXiv Machine Learning
Jul 7

Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses

arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).

By Steffen Dereich, Arnulf Jentzen, Adrian Riekert