arXiv:2604.10701v2 Announce Type: replace-cross
Abstract: Credit assignment is a central challenge in reinforcement learning (RL). Classical actor-critic methods address this challenge through fine-g...
By Zikang Shan, Han Zhong, Liwei Wang, Li Zhao
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.
By Bilal Faye, Hanane Azzag, Mustapha Lebbah
arXiv:2606. 01160v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used with formal interactive theorem provers such as Lean 4.
By Shihao Ji, Haotao Tan, Zihui Song, Mingyu Li
The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.
By Taehyung Kim, Jongeun Choi
arXiv:2402. 06359v2 Announce Type: replace Abstract: One of today's most pressing societal challenges is building AI systems whose behaviour, or the behaviour it enables within communities of interacting human and artificial agents, aligns with relevant human values.
By Nardine Osman, Mark d'Inverno
The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.
arXiv:2602. 03160v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles.
By Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do
arXiv:2510. 11194v3 Announce Type: replace Abstract: Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions.
By Peiming Li, Zhiyuan Hu, Yang Tang, Shiyu Li, Xi Chen
arXiv:2607. 15740v1 Announce Type: cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustworthy generative AI.
By Bo-An Chang, Yu-Chih Chen
arXiv:2604. 08168v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback.
By Jindi Lv, Hao Li, Jie Li, Fankun Kong, Yang Wang, Pengfei Yi, Yifei Nie, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, Guan Huang
The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.
By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
arXiv:2510. 26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems.
By Mehar Bhatia, Shravan Nayak, Gaurav Kamath, Marius Mosbach, Karolina Sta\'nczak, Vered Shwartz, Siva Reddy