arXiv AI By Yuzhou Liu, Xiyang Hu

Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning

Read the original on arXiv AI →

arXiv:2608. 01014v2 Announce Type: replace-cross Abstract: Preference optimization improves mathematical reasoning in large language models (LLMs), but reliable chosen-rejected pairs usually require verified answers, human annotations, or external reward models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

arXiv:2606. 01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO).

By Xiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu, Zhipeng Wang, Huayu Li, ZhengXiao He, Xuanzhao Dong, Prayag Tiwari, Mingkun Xu, Yujian Xiong, Feng Luo, Abolfazl Razi, Brendan Hogan Rappazzo, Anderson Schneider, Yuriy Nevmyvaka
arXiv AI
Sep 4

Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment

The paper argues that in offline preference optimization for reasoning models, applying gradients uniformly to all chosen–rejected pairs is inefficient and can be harmful. It introduces the concept of gradient utility, showing that a pair’s contribution depends on both informativeness and stability, and that high-gradient samples often lie in high‑curvature regions, causing noisy updates. To address this, the authors propose SAGE (Stability‑Aware Gradient Efficiency), which maintains difficulty‑stratified candidate pools and selects only high‑utility pairs for backpropagation, resulting in smoother optimization and better performance on mathematical reasoning benchmarks.

By Hui Wu, Hengyi Cai, Jinman Zhao, Xinran Chen, Ziheng Li, Zhejun Zhao, Shuaiqiang Wang, Yuchen Li, Dawei Yin
arXiv AI
Jun 4

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

arXiv:2606. 04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset.

By Guangcheng Zhu, Shenzhi Yang, Haobo Wang, Xing Zheng, Yingfan MA, Xuening Feng, Zhongqi Chen, Bowen Song, Weiqiang Wang, Gang Chen