arXiv Machine Learning

Generalized Score Matching for Parameter Estimation on Convex Domains

The paper introduces a generalized score matching objective for parameter estimation on convex subsets of ρ^d, derived from Minimum Probability Flow learning. It shows that this objective is a proper local scoring rule of second order, ensuring recovery of the true density when minimized, and proves convexity and consistency for exponential family models under standard conditions. Experiments demonstrate the method’s effectiveness on constrained domains where the partition function is intractable, including a generative modeling use‑case.

arXiv Statistics ML
Sep 3

Robust Bayesian Inference for Unnormalized Models with Mixed-Domain Data

The paper introduces SME-BETEL, a semiparametric Bayesian method that merges score matching estimating equations with Bayesian exponentially tilted empirical likelihood to perform inference on models with intractable normalizing constants. SME-BETEL avoids evaluating these constants and eliminates the need for learning-rate calibration, while providing consistency, asymptotic normality, and a Bernstein‑von Mises theorem that guarantees asymptotically calibrated credible sets even under model misspecification. The authors extend the framework to mixed‑domain data, enabling robust inference for doubly‑intractable models such as spatial preferential sampling, and demonstrate its effectiveness through simulations and an ozone‑monitoring application.

By Jiongran Wang, Debdeep Pati, Anirban Bhattacharya
arXiv Machine Learning
Jun 26

Learning from a Biased Sample

arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.

By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv Machine Learning
Aug 31

When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?

The paper investigates when conditional flow matching (CFM) can replace pointwise negative log-likelihood (NLL) calculations. It shows that for linear Gaussian paths, the endpoint NLL can be exactly decomposed into entropy, a weighted CFM objective, and residual terms, meaning CFM-only estimates are exact only when these residuals cancel. The study finds that ordinary CFM is generally not a pointwise NLL estimator, and even weighted variants may not fully eliminate bias, especially in training or on‑policy settings, with experiments confirming these theoretical insights.

By Yansen Han, Hongxin Sun, Tao Lin
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto