arXiv Machine Learning

Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

CluSTER is a cluster‑aware balanced sampling framework designed to improve the efficiency of instruction‑tuning for large language models. It reduces redundant computation by clustering data in gradient space and allocating samples across GPUs in a data‑parallel setting, while preserving the original distribution through weighted updates. Experiments show that CluSTER can cut training time by up to 69.6% with negligible loss in accuracy compared to existing sampling methods.

arXiv Machine Learning
Sep 25

Concurrent Split Learning Through Stable Client Clustering

The paper introduces Global Clustered Parallel Split Learning (GCPSL), a method that groups clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that distributing the population across more workloads increases direct data participation, though smaller clusters may slightly reduce accuracy. In a practical four‑GPU setup, label‑aware GCPSL achieves 85 % CIFAR‑10 validation accuracy in roughly 6 minutes, compared to over 19 minutes for serialized workloads, with size‑balanced cluster assignments improving participation by 3.25 percentage points.

By Mohammad Kohankhaki, Valentin Rentschler, Anke Schmeink
arXiv Machine Learning
Sep 4

BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training

The paper introduces BASP, a batch‑aware sequence parallelism method that partitions GPUs into disjoint groups based on micro‑batch size to reduce all‑to‑all communication. By localizing communication, BASP improves training efficiency for long‑context LLMs. Experiments on NVIDIA A100 clusters show up to 1.17‑1.31× faster end‑to‑end training on Llama and Qwen models while maintaining the same accuracy and memory usage.

By Bigyan Ghimire, Jon C. Calhoun
Hugging Face Trending Papers
Sep 24

Concurrent Split Learning Through Stable Client Clustering

The paper introduces Global Clustered Parallel Split Learning (GCPSL), which partitions clients into fixed clusters and runs Parallel Split Learning with Global Sampling (GPSL) concurrently across these clusters, periodically merging client and server model segments. Experiments with 256 logical clients show that increasing the number of concurrent workloads boosts direct data participation, though smaller clusters may slightly reduce accuracy. On a four‑GPU setup, label‑aware GCPSL achieves 85% CIFAR‑10 validation accuracy in about 6.13 minutes, compared to 19.09 minutes for serialized workloads, and size‑balanced cluster assignments improve participation by 3.25 percentage points.