arXiv:2606. 08691v1 Announce Type: new Abstract: Modern data-driven applications increasingly involve learning from multiple heterogeneous sources, where a target dataset is limited but related information is available across domains.
By Samhita Pal, Tian Gu
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.
By Diego Marcondes, Cl\'audia Peixoto
arXiv:2412. 18081v3 Announce Type: replace-cross Abstract: We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets.
By Jae Ho Chang, Massimiliano Russo, Subhadeep Paul
arXiv:2606. 18567v1 Announce Type: cross Abstract: This paper presents a methodology-centered transfer learning framework for fragility adaptation under domain shift, class imbalance, and scarce target labels while preserving engineering interpretability and supporting decision-making under uncertainty.
By Narges Saeednejad, Jamie Ellen Padgett
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.
By Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu
arXiv:2606. 27269v1 Announce Type: cross Abstract: Reliably quantifying predictive uncertainty is difficult for complex, high-dimensional, or misspecified models.
By Graham Gibson, John Tipton, Kellin Rumsey, Natalie Klein
arXiv:2603. 15158v2 Announce Type: replace Abstract: Addressing the domain adaptation problem becomes more challenging when distribution shifts across domains stem from latent confounders that affect both covariates and outcomes.
By Zahra Rahiminasab, Reza Soumi, Arto Klami, Samuel Kaski
arXiv:2607. 03005v1 Announce Type: new Abstract: In high-dimensional Ising model estimation, target sample sizes are often limited, and effectively using auxiliary binary datasets of unknown relevance remains challenging.
By Joonho Kim, Seyoung Park
arXiv:2608. 09074v1 Announce Type: cross Abstract: We develop a new approach to Personalized Federated Learning across heterogeneous clients using Nonparametric Empirical Bayes (NPEB).
By Jae Ho Chang, Arnab Auddy, Subhadeep Paul
arXiv:2607. 23404v1 Announce Type: new Abstract: Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, expensive ones, with a predictive surrogate at its core.
By Jaewook Lee, Ethan Errington, Christian D. Lorenz, Miao Guo
The paper introduces a new approach to source distribution estimation (SDE) in simulation-based science, addressing limitations of existing methods that rely on a fixed surrogate likelihood. By employing an expectation‑maximization framework, the authors iteratively train an amortized posterior on fresh simulations (E‑step) and refit the source distribution to the posterior’s average (M‑step). Two parameterizations are explored: separate source and posterior flows, and a single shared conditional flow, with experiments on three benchmark tasks showing improved performance over fixed surrogate and iterated baseline methods, notably achieving higher data‑space C2ST scores on the Lotka–Volterra benchmark.
By Trung-Dung Hoang, Lisa M. Koch