arXiv Machine Learning

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

arXiv:2608. 07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy.

arXiv Machine Learning
Sep 2

FlexP-SFT: A Flexible Aggregation-Free Framework for On-Device Personalized Split Federated Fine-Tuning of LLMs

FlexP-SFT introduces an aggregation-free framework for personalized split federated fine-tuning of large language models, eliminating the client-side aggregation step that traditionally causes communication bottlenecks and straggler issues. The method employs a layer‑flexible alignment strategy to balance personalization and generalization without global synchronization, and formulates split‑ratio selection as a resource‑aware discrete optimization problem. Experiments demonstrate that FlexP-SFT improves both accuracy and latency compared to baselines, achieving a superior resource‑accuracy trade‑off.

By Jiaxiang Geng, Tianjun Yuan, Pengchao Han, Ying Gao, Xianhao Chen, Bing Luo
arXiv Machine Learning
Sep 18

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.

By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
arXiv Computation and Language
Sep 23

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.

By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen