arXiv:2608. 08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances.
By Sai Srikar Boddupalli
arXiv:2606. 26037v1 Announce Type: cross Abstract: Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation.
By Guangzheng Hu, Patricia Men\'endez, Feng Liu, Mingming Gong, Guanghui Wang, Liuhua Peng
arXiv:2605.08992v2 Announce Type: replace
Abstract: Federated learning (FL) is increasingly used to fine-tune foundation models (FMs) on distributed private data. The community largely assumes that l...
By Kiran Naseer, Umar Shoaib
The paper investigates the trade‑off between the costs of participating in federated learning (privacy, communication, compute) and the potential gains in model performance, framing this as a game‑theoretic problem of individual rationality versus autarky. It shows that clients can remain below their local‑training baseline for many rounds and that simply capping per‑round contributions harms learning. The authors propose a new mechanism that provides short‑term participation guarantees and personalized model evaluation, demonstrating theoretically and empirically that clients can avoid short‑term losses without significantly harming overall performance, even under moderate heterogeneity.
By Amin Meghrazi, Srinivasan Parthasarathy, Andrew Perrault
arXiv:2608. 02250v1 Announce Type: new Abstract: Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models.
By Yuan-Heng Tsai, Li-Hsing Yen, Yan-Wei Chen
arXiv:2609.39250v1 Announce Type: new
Abstract: Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a...
By Muzaffer Citir, Hiroki Nishikawa, Sangyoung Park
arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.
By Xianjie Guo, Yuwei Wang, Guodu Xiang, Xiaoli Tang, Kui Yu, Han Yu, Qiang Yang
arXiv:2405. 16472v2 Announce Type: replace Abstract: Contemporary AI faces the challenge of balancing generality with user-specific personalization.
By Shutong Chen, Guodong Long, Tianyi Zhou, Jie Ma, Jing Jiang, Chengqi Zhang
arXiv:2607. 01474v1 Announce Type: new Abstract: Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard centralized techniques due to privacy and heterogeneity constraints.
By Haemin Park, Diego Klabjan, Martin W. Braun, Xiuqi Li, Balakrishnan Ananthanarayanan
The paper addresses the mismatch between learner and client data distributions in federated learning, noting that traditional client selection methods often ignore this misalignment. It introduces a dynamic, influence-aware client selection framework that uses a small proxy dataset to estimate each client's utility for the learner’s objective, prioritizing informative sources while mitigating noise and heterogeneity. Experiments on CIFAR-10 with heterogeneous partitions show the proposed method outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.
By Yiming Xie, Lili Su, Ningfang Mi
The paper introduces the first federated algorithm for Targeted Maximum Likelihood Estimation (TMLE), enabling hospitals, banks, or registries to perform TMLE without sharing individual data. Two frameworks—FedTMLE‑G, which aggregates local gradients, and FedTMLE‑L, which allows each institution to complete its own fluctuation fit—are presented, along with a finite‑precision communication protocol that keeps numerical targeting error negligible. The authors also discuss privacy implications, convergence bounds, and trade‑offs between institutional influence and sampling variability.
By Diyang Li, Fei Wang, Kyra Gan
Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and variations in dataset sizes, which can lead to biased model updates and hinder convergence.