arXiv AI

FedPS: Federated Preprocessing for structured data via aggregated Statistics

FedPS is a federated preprocessing framework that uses aggregated statistics to address missing values, inconsistent formats, and heterogeneous feature scales in structured data. It employs data-sketching techniques to summarize local datasets efficiently, enabling federated algorithms for feature scaling, encoding, discretization, and missing-value imputation. The framework also extends preprocessing-related models, such as Bayesian Linear Regression, to both horizontal and vertical federated learning settings, offering communication‑efficient and consistent pipelines for practical deployments.

arXiv Machine Learning
Aug 27

Differentiated Aggregation to Improve Generalization in Federated Learning

The paper proposes a new federated learning approach called FedALS that reduces communication costs by varying aggregation frequencies across model layers. It derives tighter generalization bounds for one‑round and multi‑round federated learning, linking these bounds to local updates and data heterogeneity. Based on representation‑learning insights, the authors argue that infrequent aggregation of early layers and more frequent aggregation of final layers yields more generalizable models, especially in non‑iid settings, and demonstrate the method’s effectiveness experimentally.

By Peyman Gholami, Hulya Seferoglu
arXiv Machine Learning
Jun 9

Federated Large Language Models: Current Progress and Future Directions

arXiv:2409. 15723v3 Announce Type: replace Abstract: Large Language Models have achieved impressive performance across diverse applications, yet their training typically depends on centralized data collection, raising serious privacy and governance concerns.

By Yuhang Yao, Jianyi Zhang, Junda Wu, Chengkai Huang, Yu Xia, Tong Yu, Ruiyi Zhang, Sungchul Kim, Ryan Rossi, Ang Li, Lina Yao, Julian McAuley, Yiran Chen, Carlee Joe-Wong
arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv AI
Sep 11

Influence-Oriented Personalized Federated Learning

Influence-Oriented Personalized Federated Learning (FedC^2I) introduces a framework that quantifies both client-level and class-level influence to enable adaptive parameter aggregation in federated learning. By modeling inter-client influence through influence vectors and matrices, FedC^2I allows clients to selectively acquire knowledge from similar peers and guides the aggregation of feature representations and classifiers. Experiments under non-IID settings show that FedC^2I outperforms existing federated learning methods in effectiveness, robustness, and interpretability.

By Yue Tan, Guodong Long, Jing Jiang, Chengqi Zhang
Hugging Face Trending Papers
Aug 10

FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning

Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and variations in dataset sizes, which can lead to biased model updates and hinder convergence.