arXiv:2608. 12108v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training across distributed clients while keeping data local.
By Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay
ConsensusBench is a new dataset that supplies rule‑based process‑level signals for large language model reasoning. It identifies key intermediate conclusions—called Consensus Nodes—by filtering correct trajectories and clustering semantically equivalent statements. By incorporating a process reward derived from these nodes into GRPO‑style reinforcement learning, the authors create ConsensusPR, which reduces reward sparsity and improves performance on benchmarks such as AIME, GSM8K, and MATH‑500.
By Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling
arXiv:2609.00213v1 Announce Type: new
Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific e...
By Yu Yuan, Yaoyou Fan, Lili Zhao, Guangting Zheng, Kai Zhang, Lu Pan, Ke Zeng, Qi Liu
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments.
arXiv:2606. 15625v1 Announce Type: new Abstract: The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation.
By Yijun Lu, Zihan Fang, Pengpeng Qiao, Zheng Lin, Jing Yang, Yuxin Zhang, Por Lip Yee, Zhe Chen, Jun Luo
arXiv:2608. 01556v1 Announce Type: new Abstract: Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized.
By Seongyoon Kim, Boryeong Cho, Jihwan Oh, Seokhyun Chung, Se-Young Yun
arXiv:2608. 10499v1 Announce Type: cross Abstract: Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy.
By Md Rafid Islam, Rafsan Jany, Zahid Hasan, Ratun Rahman
arXiv:2609.07312v1 Announce Type: new
Abstract: This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attack...
By Xiao Ma, Hong Shen, Hui Tian, Wenqi Lyu, Wei Ke
arXiv:2604.16778v3 Announce Type: replace-cross
Abstract: Modern agents specialize in varying domains while there is no clear approach combining different domain skills. We propose a federated learni...
By Dixi Yao, Tahseen Rabbani, Manzil Zaheer, Tian Li
arXiv:2606. 07950v1 Announce Type: new Abstract: RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute allocation.
By Zhanke Zhou, Xiangyu Lu, Chentao Cao, Brando Miranda, Tongliang Liu, Bo Han, Sanmi Koyejo
The paper introduces a latent information sharing scheme for federated learning that mitigates client drift by sharing a small amount of hidden‑layer activations. The authors demonstrate both theoretically and empirically that this approach improves training efficiency while maintaining convergence guarantees and data privacy. Compared to existing methods such as FedProx, SCAFFOLD, FedPVR, FedProto, and SplitFed, the proposed method achieves higher model accuracy within a fixed round budget without adding significant communication overhead.
By Seungjun Lee, Ensieh Khazaei, Dimitrios Hatzinakos, Baturalp Buyukates, Sunwoo Lee
GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models proposes a new method for fine‑tuning LLMs after training. The approach models each group gradient as a random variable, estimates its probability distribution, and uses Dirichlet‑based gradient uncertainty to weight each group’s contribution during policy updates. Experiments on multiple benchmarks show that this uncertainty‑aware aggregation improves the effectiveness of post‑training policy optimization.
By Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan, Jiahuan Zhou, Changwen Zheng, Wenwen Qiang