arXiv Computation and Language

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

HARPO is a reinforcement learning framework that jointly optimizes faithfulness and creativity in language generation. It uses a Hallucination-Aware Generative Reward Model (HA‑GRM) to evaluate both faithfulness and writing quality, and a Selective Activation Mechanism (SAM) that applies writing rewards only to hallucination‑free outputs. Experiments on Qwen models show that HARPO improves faithfulness scores and reduces hallucination rates while boosting creative‑writing performance.

arXiv AI
Jun 10

TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

arXiv:2509. 25760v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge.

By Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang, Rulin Shao, Jingxiang Chen, Mohammad Kachuee, Teja Gollapudi, Yiwei Liao, Nicolas Scheffer, Rakesh Wanga, Anuj Kumar, Yu Meng, Wen-tau Yih, Xin Luna Dong
arXiv AI
6d ago

What Pretraining and Midtraining Make Learnable from Rewards?

The paper investigates how pretraining and midtraining enable reward-based learning by providing necessary information and computation. It analyzes sequential state computation and contextual memory, showing that task‑independent source observations resolve ambiguities in reward adaptation. Experiments on pretrained Qwen2.5 checkpoints across eight worlds demonstrate that correct source and first‑operation supervision significantly improve success rates, and that memory replay and independent confirmation further enhance performance.

By Chiwun Yang, Xiaoyu Li
arXiv Computation and Language
Sep 4

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

The paper introduces CreaEval, an automated creativity evaluator designed for complex multi-step tasks (CGPST). It separates the evaluation process into two phases: a memory‑augmented analysis that transforms responses into structured evidence, and an evidence‑based judging step that scores without seeing raw outputs. Experiments show CreaEval outperforms existing baselines by an average of 22.74% across CGPST and two simpler creativity tasks.

By Xiangyu Wang, Jin Wu, Xiaoyu Li, Chanjin Zheng, Yifeng Zhou
arXiv AI
Jul 14

To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning

arXiv:2607. 10738v1 Announce Type: cross Abstract: Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks.

By Fengji Zhang, Tianyu Fan, Yuxiang Zheng, Xinyao Niu, Chengen Huang, Jacky Keung, Bei Chen