arXiv AI By Jinkun Hou, Zhuo Liu, Huimin Ren, Hongsheng Xin, Pan Zhou, Kun Zhan

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

Read the original on arXiv AI →

arXiv:2608. 09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 30

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation.

arXiv Computation and Language
Sep 1

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

PaperGym is a framework that transforms each research paper into a training environment for AI research planning, using rubrics extracted from the paper’s method and experiments as a critic. It synthesizes research questions from the goal and background, and derives evaluation criteria from the method and experiments, reducing criterion leakage to 3.7%. Experiments with Qwen models show that training with PaperGym’s rubric improves benchmark performance by up to 5.6 points and outperforms existing datasets and fine‑tuning baselines.

By Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen