arXiv Machine Learning By Wenyang Hu, Junxiang Jia, Zhen Shu, Daniel Dahlmeier, See-Kiong Ng, Bryan Kian Hsiang Low

ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 24994v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.