arXiv AI

FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment

arXiv:2606. 01607v1 Announce Type: cross Abstract: Federated learning (FL) is a decentralized approach that enables collaborative model training without exposing raw data.

arXiv Machine Learning
Aug 10

FedTransKD-IDS: Robust Federated Transfer Learning with Knowledge Distillation for Intrusion Detection in IoT

arXiv:2608. 06447v1 Announce Type: cross Abstract: In modern distributed network environments, particularly in Internet of Things infrastructures and 5G networks, stringent privacy preservation and scalability requirements have created significant challenges for intrusion detection systems.

By Mohammad Hosssein Gholamrezazadeh, Ahmadreza MontazerolghaemAhmadreza Montazerolghaem
arXiv Machine Learning
5d ago

Distributed Learning as a Service: The Developer's Perspective

The paper introduces Distributed Learning as a Service (DLaaS), a platform that lets developers launch distributed/federated learning jobs through a single admin dashboard. It offers declarative options such as Differential Privacy, Split Learning, Hierarchical Aggregation, and Knowledge Distillation without requiring changes to client code. The authors demonstrate the full service lifecycle on an industrial smart‑home Wake‑up Word task using the Ok Aura dataset, showing live operation across Android clients and Dockerized aggregators, and releasing the source code and video walkthroughs.

By Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando L\'opez, Pablo Gomez Guerrero, Sotirios Spantideas, David Solans Noguero
arXiv AI
Sep 18

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.

By Gianluca Mittone, Marco Aldinucci