arXiv Machine Learning

Formal Bayesian Transfer Learning via the Total Risk Prior

arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto
arXiv Machine Learning
Jun 18

Bridging Data Gaps in Structural Fragility Modeling through Transfer Learning: Methodology and Case Studies

arXiv:2606. 18567v1 Announce Type: cross Abstract: This paper presents a methodology-centered transfer learning framework for fragility adaptation under domain shift, class imbalance, and scarce target labels while preserving engineering interpretability and supporting decision-making under uncertainty.

By Narges Saeednejad, Jamie Ellen Padgett
arXiv Machine Learning
Aug 12

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.

By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu
arXiv Machine Learning
Sep 3

Source Distribution Estimation by Posterior Averaging

The paper introduces a new approach to source distribution estimation (SDE) in simulation-based science, addressing limitations of existing methods that rely on a fixed surrogate likelihood. By employing an expectation‑maximization framework, the authors iteratively train an amortized posterior on fresh simulations (E‑step) and refit the source distribution to the posterior’s average (M‑step). Two parameterizations are explored: separate source and posterior flows, and a single shared conditional flow, with experiments on three benchmark tasks showing improved performance over fixed surrogate and iterated baseline methods, notably achieving higher data‑space C2ST scores on the Lotka–Volterra benchmark.

By Trung-Dung Hoang, Lisa M. Koch