arXiv AI

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

arXiv:2607. 28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets.

arXiv Machine Learning
Jul 24

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

arXiv:2607. 20465v1 Announce Type: new Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end.

By Hao Liang, Qifeng Cai, Yibo Lin, Jianzhuo Du, Qifeng Xia, Sizhe Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, Bohan Zeng, Ruichuan An, Conghui He, Wentao Zhang
arXiv AI
Sep 15

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

FLoKD is an adaptive knowledge‑distillation framework designed for federated fine‑tuning of low‑rank LLMs over wireless networks. It transmits intermediate LoRA activations instead of full parameters or token‑level logits, and uses transformer block importance scoring plus dataset selection to reduce communication. Experiments on WikiText‑103, PTB, and Dialog show a 50‑65% reduction in communication while maintaining competitive perplexity.

By Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi
arXiv Machine Learning
Sep 22

Joint Domain-Class Modeling for Federated Learning Under Feature Skew

The paper introduces Joint Domain-Class Federated Learning (JDFL), a lightweight, optimizer‑agnostic extension designed to address feature skew in federated learning. JDFL infers pseudo‑domains from local update signals and expands the classifier head to output joint domain‑class logits, enabling the model to capture domain‑conditioned appearance while sharing a backbone. Two supervision strategies—similarity‑aware soft‑labeling and per‑sample randomized target assignment—are proposed to train the expanded head, and experiments on domain‑shifted image benchmarks show consistent improvements in global test accuracy over standard FL methods.

By Sina Najafi, Mostafa Tavassolipour, Seyed Pooya Shariatpanahi