arXiv Machine Learning By Jinliang Gao, Ning Yang, Hai Wang, Baili Xiao, Pin Lyu

SDO: Structure-Aware Data Organization for Efficient LLM Post-Training

Read the original on arXiv Machine Learning →

arXiv:2607. 27273v1 Announce Type: new Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
3d ago

DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models

arXiv:2603.26164v2 Announce Type: replace-cross Abstract: Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters...

By Hao Liang, Zhengyang Zhao, Mingrui Chen, Meiyi Qiang, Lu Ma, Rongyi Yu, Hengyi Feng, Shixuan Sun, Zimo Meng, Xiaochen Ma, Xuanlin Yang, Qifeng Cai, Ruichuan An, Bohan Zeng, Zhen Hao Wong, Chengyu Shen, Runming He, Zhaoyang Han, Yaowei Zheng, Fangcheng Fu, Conghui He, Bin Cui, Zhiyu Li, Weinan E, Wentao Zhang
arXiv Machine Learning
Sep 14

Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

CluSTER is a cluster‑aware balanced sampling framework designed to improve the efficiency of instruction‑tuning for large language models. It reduces redundant computation by clustering data in gradient space and allocating samples across GPUs in a data‑parallel setting, while preserving the original distribution through weighted updates. Experiments show that CluSTER can cut training time by up to 69.6% with negligible loss in accuracy compared to existing sampling methods.

By Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee