arXiv AI By Mengchen Li

AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution

Read the original on arXiv AI →

arXiv:2607. 08252v1 Announce Type: new Abstract: Long-term persona agents must remain identifiable while adapting to new events, relationships, evidence, and social conditions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 31

Living-Harness Is an Interactive-Agent Evolver

arXiv:2607. 26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions.

By Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu
arXiv AI
Sep 7

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

The paper introduces an online skill‑evolution framework that transforms interaction traces and evaluator feedback into a persistent, versioned library of reusable procedures for computer‑use agents. By executing each iteration against a frozen library snapshot, the system updates skills without altering the underlying model parameters. Experiments across four OSWorld domains show that the evolving library consistently outperforms an empty‑library baseline, with gains ranging from 5.7 to 18.6 percentage points, while also revealing domain‑specific temporal stability and challenges in skill retrieval and revision.

By Longtao Hu, Xiao Liang, Linchao Zhu
arXiv Computation and Language
Aug 28

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

ContextEcho is a benchmark and harness designed to measure persona drift in large language models during long, tool‑using coding sessions. It includes a 25‑probe identity suite, a snapshot‑then‑probe protocol that preserves the main conversation, and both judged and judge‑free measurement surfaces. Across 23 frontier models and thousands of turns, the benchmark shows that persona drift is widespread, not limited to specific model families, and that simple in‑session compaction does not reset it, while a single‑shot anchor can restore the intended persona.

By Xianzhong Ding, Yangyang Yu, Changwei Liu, Bill Zhao, Le Chen, Tao Chen