arXiv AI

Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction

arXiv:2607. 18973v1 Announce Type: cross Abstract: Textual skills provide a lightweight way to improve frozen language-model agents, but their self-evolution normally requires a stable validation signal.

arXiv Computation and Language
Sep 4

RL-ADA: A World-Feedback Framework for Adversarially Robust Enterprise Dialogue Agents

RL-ADA introduces a co‑evolutionary training framework that replaces costly human annotations with world‑feedback rewards derived from interaction outcomes. In this system, a large Customer Support Agent and an Adversarial Customer Agent train together, guided by an automated judge that rewards successful resolution and realistic intent‑concealing utterances, respectively. Applied to a banking support proof of concept, the method eliminates routing errors and doubles the end‑to‑end PASS rate over five cycles, while also revealing a new adversarial strategy called Contextual Camouflage.

By Ram Narayanan, Harshit Rajgarhia, Abhishek Mukherji
arXiv Computation and Language
Sep 3

PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation

PragAlign is a feedback‑guided framework that improves synthetic dialogue generation by iteratively generating, evaluating, and revising conversations to meet specified service context, target intent, and target emotion. Using an LLM‑based evaluator that scores intent alignment, emotion alignment, coherence, fluency, and overall quality, PragAlign achieves a 99.50% acceptance rate on 800 dialogue specifications, outperforming one‑shot and repeated generation without feedback. Human evaluation confirms that intent expression and dialogue flow are reliably recognized, while emotion appropriateness remains more variable.

By Smitha Muthya Sudheendra, Jaideep Srivastava
arXiv AI
Sep 25

J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data

J-Zero introduces a unified Challenger–Solver–Judge self‑evolution framework that operates without any initial data. The Challenger and Solver co‑evolve through adversarial task generation and response improvement, while the Judge adapts using known preference pairs derived from the Solver’s outputs rather than its own scores. Experiments show J‑Zero surpasses baselines by 4.2 points on verifiable tasks and 8.0 points on unverifiable tasks, maintaining improvement over ten iterations versus baseline degradation after two.

By Gyouk Chu, Myeongho Jeon, Teresa Yeo, Eunho Yang
arXiv AI
Aug 28

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

J-Zero introduces a unified Challenger–Solver–Judge co‑evolution framework that enables self‑improvement of language models without requiring external supervision. The Challenger generates increasingly difficult tasks, the Solver learns to produce better responses, and the Judge adapts using preference pairs derived from the Solver’s own outputs rather than from its own scores. Experiments show J‑Zero outperforms baselines by an average of 4.2 points on verifiable tasks and 8.0 points on unverifiable tasks, maintaining improvement over ten iterations while baselines degrade after two.

By Gyouk Chu, Myeongho Jeon, Eunho Yang
arXiv Machine Learning
Sep 15

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

arXiv:2609.15972v1 Announce Type: cross Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...

By Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
arXiv AI
Jul 21

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.

By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang