arXiv:1907.06994v2 Announce Type: replace-cross
Abstract: Mixtures of experts (MoE) are conditional mixture models in which both the mixing proportions and the component densities depend on the predi...
By Thin Nguyen-Van, Faicel Chamroukhi, Ha Hoang Van, Bao Tuyen Huynh
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
The paper introduces an individualized sparse regression framework for matrix‑valued covariates, where each observation has its own relevant rows while regression effects are shared across the population. It proposes a diagonalized attention mechanism that uses query–key scores to localize sample‑specific signal rows and a value matrix for downstream regression, achieving a parameter dimension independent of sample size. The authors provide existence theorems guaranteeing recovery of latent rows under score‑separation and concentration conditions, and demonstrate strong prediction, localization, and classification performance in simulations and real sentiment analysis.
By Borui Peng, Liwei Lin, Feifei Wang, Long Feng
arXiv:2607. 02681v1 Announce Type: cross Abstract: Integrating information across related tasks can improve estimation and prediction in transfer, multi-task, and federated learning, but contamination and heterogeneity make robust borrowing challenging.
By Ye Tian, Mengchu Li, Marco Avella Medina
The paper introduces COVER, a multi‑task learning framework that regularizes covariate overlap to mitigate the negative effects of sharing information across tasks with differing covariate distributions and response relationships. COVER blends a common component function, a shared neural representation, and low‑dimensional task‑specific coefficients, using taskwise second‑moment matrices to guide coefficient integration. The authors provide theoretical bias‑variance analysis, oracle inequalities, and neural‑network convergence rates, and demonstrate that COVER outperforms existing deep‑learning and statistical integration methods in simulations and a GTEx central‑nervous‑system study.
By Yang Sui, Qi Xu, Yang Bai, Annie Qu
arXiv:2609.22654v1 Announce Type: cross
Abstract: Federated learning (FL) has emerged as a leading privacy-preserving framework for collaborative machine learning across decentralized environments. W...
By Brigham Halverson, Sharmistha Guha, Jessica Bernard, Rajarshi Guhaniyogi
FedFIbOS introduces a Fisher‑importance based criterion for selecting submodel parameters in heterogeneous federated learning, addressing the lack of theoretical justification in prior heuristic methods. By deriving a Fisher‑weighted quadratic masking surrogate and showing that the raw Fisher top‑k rule satisfies this surrogate under a Fisher‑dominant ranking condition, the method preserves convergence guarantees while efficiently estimating Fisher scores from squared gradients. Experiments on CIFAR‑10, CIFAR‑100, and AGNews demonstrate that FedFIbOS outperforms state‑of‑the‑art approaches by roughly 10% in accuracy, especially under strong non‑IID heterogeneity.
By Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang
arXiv:2607. 22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset.
By L{\'e}o Nicollier (CB, ATT), Marc Pic (ATT), Pablo Mus{\'e} (CB, IFUMI), Enric Meinhardt-Llopis (CB), Gabriele Facciolo (CB)
The paper introduces a sample-weighted trace-norm geometry for multitask learning, focusing on the end-to-end map from task coefficients to input-space predictors. It derives the exact empirical Rademacher complexity for a fixed-radius class and demonstrates that this measure captures orientation and factorization effects that traditional product-based bounds miss. Experiments on 252 held-out comparisons show that weighted joint nuclear regularization consistently improves population excess error over unweighted nuclear regularization and other baselines.
By Mahdi Mohammadigohari
arXiv:2607. 04085v1 Announce Type: cross Abstract: Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level intelligence: at inference time, the server routes each external query to the best-matched client for prediction.
By Zijian Wang, Pengfei Li, Guangyu Yang, Qiong Zhang
CORE-STACK+ is a new meta‑learning framework for deep stacked generalization that tackles two key problems in heterogeneous vision ensembles: prediction‑space multicollinearity and calibration collapse. It introduces a four‑step preconditioning pipeline—kernelized redundancy filtering, a lightweight differentiable meta‑feature gate, a spectrum‑adaptive ridge penalty, and a Laplace‑approximate Bayesian blender—to jointly improve conditioning and calibration. Across six vision benchmarks, CORE‑STACK+ boosts accuracy, reduces model count and inference cost, and significantly lowers expected calibration error compared to existing methods.
By Noor Islam S. Mohammad
arXiv:2605. 27991v2 Announce Type: replace-cross Abstract: Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early-stopping rules.
By Minhao Yao, Ruoyu Wang, Xihong Lin, Lin Liu, Zhonghua Liu