arXiv Computation and Language By Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

Read the original on arXiv Computation and Language →

HARPO is a reinforcement learning framework that jointly optimizes faithfulness and creativity in language generation. It uses a Hallucination-Aware Generative Reward Model (HA‑GRM) to evaluate both faithfulness and writing quality, and a Selective Activation Mechanism (SAM) that applies writing rewards only to hallucination‑free outputs. Experiments on Qwen models show that HARPO improves faithfulness scores and reduces hallucination rates while boosting creative‑writing performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 10

TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

arXiv:2509. 25760v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge.

By Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang, Rulin Shao, Jingxiang Chen, Mohammad Kachuee, Teja Gollapudi, Yiwei Liao, Nicolas Scheffer, Rakesh Wanga, Anuj Kumar, Yu Meng, Wen-tau Yih, Xin Luna Dong
arXiv AI
6d ago

What Pretraining and Midtraining Make Learnable from Rewards?

The paper investigates how pretraining and midtraining enable reward-based learning by providing necessary information and computation. It analyzes sequential state computation and contextual memory, showing that task‑independent source observations resolve ambiguities in reward adaptation. Experiments on pretrained Qwen2.5 checkpoints across eight worlds demonstrate that correct source and first‑operation supervision significantly improve success rates, and that memory replay and independent confirmation further enhance performance.

By Chiwun Yang, Xiaoyu Li