arXiv Machine Learning

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

arXiv:2608. 01559v1 Announce Type: cross Abstract: Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack.

arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu
arXiv Machine Learning
Aug 19

Debate Training Reduces Reward Hacking in RLAIF

The paper shows that fine‑tuning a large language model (LLM) with a debate framework—where a generator and a critic compete and a weaker LLM judge adjudicates—reduces reward hacking compared to standard reinforcement learning from AI feedback (RLAIF). In experiments on mathematics tasks, the debate approach keeps the judge’s performance stable, achieving a 45% higher peak validation accuracy than the RLAIF baseline and mitigating the rapid exploitation of judge errors. Additional findings indicate that weakening the judge speeds hacking unless countered by extra debate rounds, that debate can override misalignment prompts, and that word‑limit constraints on critiques help balance the game and prevent judge hacking. whyItMatters":"The study demonstrates a practical method to curb reward hacking in RL‑based AI systems, addressing a key obstacle for safely scaling AI oversight."

By Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
arXiv AI
Jul 24

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

arXiv:2607. 20527v1 Announce Type: new Abstract: Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their benchmarks already check whether those citations hold, with a fixed attribution model or human graders.

By Taewan Goo, Junsik Kim, Kyulhee Han, GwonYul Jo, Jong-Soo Kim, Tae-Hyung Kim
arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
arXiv AI
Sep 11

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

The paper introduces the concept of proof‑carrying cognition, aiming to close the verification gap in language‑model reasoning by using reality‑settled rewards. It presents a theoretical framework linking verifier‑gold correlation to compute‑capability trade‑offs, demonstrates that unsound verifiers degrade under best‑of‑N selection while sound verifiers improve, and proposes a new benchmark metric, Soundness‑under‑Pressure, for evaluating reality‑settled reasoning systems.

By Eshwar Reddy M, Sourav Karmakar
arXiv Machine Learning
Sep 16

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

The paper introduces ImpossibleRubrics, a benchmark of 169 impossible tasks designed to test the robustness of language‑model‑generated rubrics as reward signals. Each task is paired with a verifiable oracle certificate that defines what constitutes an honest answer, and the benchmark includes 48 answerable controls. Experiments show that many rubric generators are exploited frequently—up to 36% on a stress cut—highlighting a significant gap in rubric quality rather than task difficulty, and that generic rubrics can be more vulnerable than tailored ones.

By Bowen Qin, Yi Xie, Yesheng Liu, Xi Yang