arXiv AI

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.

arXiv AI
Jul 24

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

arXiv:2602. 10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors.

By Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
arXiv Computation and Language
Sep 1

Aspire: Can Models Self-Evolve from Vague Goals?

The paper introduces ASPIRE, a benchmark that challenges language model agents to self‑evolve from vague, natural‑language goals without explicit evaluation metrics. In ASPIRE, agents must interpret the goal, select data and update strategies, and decide when to evaluate, all while the downstream tasks remain hidden. Experiments show that while agents can complete training loops, weight‑level improvements are sparse and unstable, and the best evolved harness still falls short of a strong engineered baseline.

By Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
arXiv AI
Sep 10

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

The paper reports the first Turing test for speech‑to‑speech systems, gathering 2,968 human judgments on conversations between nine state‑of‑the‑art S2S systems and 28 humans. None of the evaluated systems passed the test, highlighting a clear gap in human‑likeness. The authors diagnose the failure with an 18‑dimension taxonomy, finding that paralinguistic cues, emotional expressivity, and conversational persona—not semantic understanding—are the main bottlenecks, and they propose an interpretable model for automatic human‑vs‑machine discrimination.

By Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang
arXiv AI
Aug 28

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

The paper investigates Self‑Generated Text Recognition (SGTR), the ability of large language models (LLMs) to identify their own outputs. By evaluating 13–21 models across 6 experimental designs, it shows that SGTR accuracy varies with evaluation format, conversation structure, and task domain, and that a quality‑heuristic bias dominates results. The study also finds that fine‑tuning for SGTR in one setting can generalize to others and may cause models to prefer their own outputs when judging, highlighting potential safety concerns.

By Jesse St. Amand, Callum Canavan, Sohaib Imran, Joseph Hewson, Aaron Lutz, Shi Feng, Puria Radmard, Lennie Wells
arXiv Computation and Language
Sep 4

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

RL-ADA introduces a co‑evolutionary training framework that replaces costly human annotations with world‑feedback rewards derived from interaction outcomes. In this system, a large Customer Support Agent and an Adversarial Customer Agent train together, guided by an automated judge that rewards successful resolution and realistic intent‑concealing utterances, respectively. Applied to a banking support proof of concept, the method eliminates routing errors and doubles the end‑to‑end PASS rate over five cycles, while also revealing a new adversarial strategy called Contextual Camouflage.

By Ram Narayanan, Harshit Rajgarhia, Abhishek Mukherji