arXiv Machine Learning

Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension

arXiv:2608. 11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data.

arXiv Machine Learning
Sep 3

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

The paper introduces soft‑label‑based estimators for the Bayes‑optimal balanced error rate (BER) and area under the ROC curve (AUC), extending from a clean setting with known class priors to a realistic scenario with unknown priors and corrupted soft labels. It also adapts the FeeBee evaluation framework to assess these estimators without needing the true optimum, providing practical evaluation scores for any estimator of optimal BER or AUC. Experiments on synthetic and real datasets confirm the effectiveness of both the estimators and the evaluation method.

By Ryota Ushio, Takashi Ishida, Masashi Sugiyama
arXiv Machine Learning
Sep 18

FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning

FedFIbOS introduces a Fisher‑importance based criterion for selecting submodel parameters in heterogeneous federated learning, addressing the lack of theoretical justification in prior heuristic methods. By deriving a Fisher‑weighted quadratic masking surrogate and showing that the raw Fisher top‑k rule satisfies this surrogate under a Fisher‑dominant ranking condition, the method preserves convergence guarantees while efficiently estimating Fisher scores from squared gradients. Experiments on CIFAR‑10, CIFAR‑100, and AGNews demonstrate that FedFIbOS outperforms state‑of‑the‑art approaches by roughly 10% in accuracy, especially under strong non‑IID heterogeneity.

By Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang
arXiv Machine Learning
Jun 26

Learning from a Biased Sample

arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.

By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv Machine Learning
Jun 19

Variational Consensus Monte Carlo for Bayesian Mixture

arXiv:2606. 19643v1 Announce Type: cross Abstract: Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.

By Julie Fendler, Francesca L. Crowe, Tom Marshall, Sylvia Richardson, Paul D. W. Kirk
arXiv Machine Learning
Sep 10

Large Classification-Risk-Optional Label Acquisition

arXiv:2609.06873v1 Announce Type: cross Abstract: We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification...

By F. Setoudehtanzangi, Geoffrey J. McLachlan