arXiv Machine Learning

Learning to bin: differentiable and Bayesian optimization for multi-dimensional discriminants in high-energy physics

The paper introduces a method for optimizing bin boundaries in multi-dimensional discriminants using a Gaussian Mixture Model (GMM), allowing flexible definition of analysis categories. Two optimization strategies—differentiable and Bayesian—are compared in toy binary and three-class setups, with the differentiable approach excelling in multi-dimensional cases. Applied to the FAIR Universe $H ightarrow au au$ dataset, the GMM-based optimization achieves the highest signal significance, and the tools are released as lightweight Python plugins.

arXiv Machine Learning
Aug 27

Multi-output Gaussian process prediction of physical fields under linear equality constraints

The paper tackles the challenge of predicting multiple high‑dimensional physical fields that must satisfy linear equality constraints, a common scenario in physics‑informed machine learning. It critiques the conventional approach of deducing one field from others, showing its sensitivity to arbitrary choices and its impact on accuracy and uncertainty. To address this, the authors introduce a symmetric framework that first applies a row‑wise PCA to preserve constraints in a latent space, then trains a linearly‑constrained multi‑output Gaussian process using a specially parametrized kernel, and validate the method on population dynamics and CFD problems involving Reynolds stress tensors.

By Mahamat Hamdan Nassouradine, Cl\'ement Gauchy, Pierre-Emmanuel Angeli, S\'ebastien da Veiga
arXiv Machine Learning
Aug 19

Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification

The paper presents an exact constrained reformulation for direct metric optimization (DMO) in binary imbalanced classification, focusing on precision, recall, and F1-score under three settings: fixing precision to optimize recall, fixing recall to optimize precision, and optimizing F1-score. Unlike prior approaches that use smooth approximations, the authors introduce exact penalty methods to solve these problems efficiently. Experiments on benchmark datasets show that this exact reformulation and optimization (ERO) framework outperforms state‑of‑the‑art methods for all three DMO tasks.

By Le Peng, Yash Travadi, Chuan He, Ying Cui, Ju Sun
arXiv Machine Learning
Aug 24

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

The paper introduces an amortized learning framework for selecting bandwidths in kernel density estimation by optimizing the logarithmic score across a distribution of tasks. It uses a truncated-and-renormalized bounded-support formulation and affine standardization to achieve stable learning and transferability across different intervals. Experiments on Gaussian samples, a multi-family benchmark, and randomized Gaussian mixtures demonstrate that the learned selector outperforms traditional methods such as Silverman’s rule, Sheather–Jones, and least‑squares cross‑validation, especially for small or heterogeneous samples.

By Junyi Liang, Hailiang Du
arXiv Machine Learning
Aug 12

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.

By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu