arXiv AI

Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes

Augur is a synthetic decision laboratory that simulates how users will react to product and policy changes before they are released. It constructs a typed knowledge graph from change documents, populates a persona market, runs simulations, and produces an auditable decision memo recommending one of five actions. Using a dataset of 50 real episodes (Gold‑50), the authors evaluate the system’s five‑way release verdicts and find that evaluation design, rather than model capability, largely drives performance differences among frontier and open‑weight models.

arXiv AI
2d ago

Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

The paper introduces TrustSwap, a counterfactual test that swaps or removes source reliability labels while keeping evidence text constant, to evaluate how retrieval‑augmented fact‑checking models respond across verdict, confidence, and search decisions. Experiments on untrained and RL‑trained models show that confidence and search largely follow labels, yet label changes can flip a significant portion of verdicts, especially in larger models. The authors propose trust‑swap augmentation (TSA) to mitigate this shortcut, demonstrating reduced verdict flip rates and maintained accuracy in several settings, though its effectiveness diminishes at larger model scales.

By Jianchang Su, Yiwei Yang, Wei Zhang
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv Machine Learning
Sep 23

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.

By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng
arXiv AI
Jun 11

Search Discipline for Long-Horizon Research Agents

arXiv:2606. 11522v1 Announce Type: new Abstract: Autoresearch agents now propose, evaluate, and select scientific candidates against a metric, and that metric is usually an aggregate reduced over a heterogeneous space of regions, slices, or cohorts.

By Adithya Srinivasan, Devesh Paragiri