arXiv Machine Learning By Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink

Concurrent Split Learning Through Stable Client Clustering

Read the original on arXiv Machine Learning →

The paper introduces Global Clustered Parallel Split Learning (GCPSL), a method that groups clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that distributing the population across more workloads increases direct data participation, though smaller clusters may slightly reduce accuracy. In a practical four‑GPU setup, label‑aware GCPSL achieves 85 % CIFAR‑10 validation accuracy in roughly 6 minutes, compared to over 19 minutes for serialized workloads, with size‑balanced cluster assignments improving participation by 3.25 percentage points.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 24

Concurrent Split Learning Through Stable Client Clustering

The paper introduces Global Clustered Parallel Split Learning (GCPSL), which partitions clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that increasing the number of concurrent workloads boosts direct data participation, though smaller clusters may slightly reduce accuracy. On a four‑GPU setup, label‑aware GCPSL achieves 85% CIFAR‑10 validation accuracy in about 6.13 minutes, compared to 19.09 minutes for serialized workloads, and size‑balanced cluster assignments improve participation by 3.25 percentage points.

arXiv Machine Learning
Sep 14

Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

CluSTER is a cluster‑aware balanced sampling framework designed to improve the efficiency of instruction‑tuning for large language models. It reduces redundant computation by clustering data in gradient space and allocating samples across GPUs in a data‑parallel setting, while preserving the original distribution through weighted updates. Experiments show that CluSTER can cut training time by up to 69.6% with negligible loss in accuracy compared to existing sampling methods.

By Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee
arXiv AI
Sep 18

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.

By Gianluca Mittone, Marco Aldinucci