arXiv AI By Jian Hong, Chen Cheng, Quan Liu, Yuhao Chen, Enhong Chen

STAIF: A Stage-wise Optimization for Complex Instruction Following

Read the original on arXiv AI →

arXiv:2607. 22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 9

TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles

arXiv:2606. 07520v1 Announce Type: cross Abstract: Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.

By Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Wu Ning, Haonan Song, Dandan Tu, Qixun Zhang, Yuxiang He, Bibo Cai, Ting Liu
Hugging Face Trending Papers
Jun 3

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization

Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground truth, creating a high-variance, subjective reward landscape. Previous preference-based fine-tuning methods have major trade-offs: Direct Preference Optimization (DPO) is limited by the lack of exploration inherent in offline training, while Proximal Policy Optimization (PPO) can lead to training instability due to potentially unreliable critic estimates.