arXiv AI By Xubo Liu, Wenya Guo, Ruxue Yan, Xinying Qian, Ying Zhang

Rewarding Better Thinking for LLM Preference Alignment

Read the original on arXiv AI →

arXiv:2607. 19824v1 Announce Type: new Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 30

To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregation

arXiv:2606. 29481v1 Announce Type: cross Abstract: While reinforcement learning (RL) significantly enhances LLM reasoning, its efficacy is severely undermined by Pre-RL data overlap, where RL datasets overlap with pretraining or SFT corpora, causing models to exploit shortcuts by memorizing correct answers and fabricating post-hoc reasoning.

By Jiuheng Lin, Chen Zhang, Yansong Feng