arXiv Machine Learning By Nan Li

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

Read the original on arXiv Machine Learning →

The paper presents a post‑training recipe for small dialogue‑game agents that involves three steps: acquiring broad game participation via supervised fine‑tuning, repairing specific local failures with turn‑local preference pairs, and preserving general capabilities. Applied to the LM Playschool Challenge, the method raises the public clemscore from 10.67 to 38.92 and the closed in‑domain score from 13.41 to 41.17 while keeping overall static performance nearly unchanged. The gains are mainly within the targeted game family, with limited improvement on out‑of‑domain clemscore.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
1d ago

First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents

The paper introduces Qwen-GuidePlay-2B, a 2B‑parameter language model fine‑tuned for dialogue‑game interaction. The training pipeline consists of three stages: supervised fine‑tuning on successful game trajectories, weighted turn‑level fine‑tuning, and teacher‑guided fine‑tuning that corrects formatting and evaluates examples. The resulting model achieves a clemscore of 57.12 and a statscore of 42.68 on Playpen’s public validation, ranking second in the official challenge and demonstrating that careful curation can outperform more aggressive procedural methods.

By Syed Mahbubul Huq, Pranava Madhyastha
arXiv AI
Jul 21

Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.

By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
arXiv AI
Aug 24

Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

The paper introduces Evaluation-as-Search (EaS), a feedback‑driven method that adaptively probes LLM‑powered meeting assistants by focusing on natural questions likely to reveal grounding failures. Using EaS, the authors build MeetingProbe, a benchmark of over 3,000 annotated question‑answer pairs from 20 transcripts across three meeting genres and three assistants. Ablation studies show that adaptive search uncovers 2.5× more failures than random probing, revealing a capability gradient and eight recurring failure categories dominated by discourse‑pragmatic challenges.

By Sami Khairy, Yasaman Hosseinkashi, Vishak Gopal, Ross Cutler
arXiv AI
Jul 14

Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points

arXiv:2607. 11363v1 Announce Type: cross Abstract: Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models' functional ToM abilities in ways that generalize to naturalistic settings.

By Roberta Rocca, Sami Boukortt, Geoff Keeling, Winnie Street