arXiv AI

J-Zero: Unified Challenger--Solver--Judge Self-Evolution from Zero Data

J-Zero introduces a unified Challenger–Solver–Judge self‑evolution framework that operates without any initial data. The Challenger and Solver co‑evolve through adversarial task generation and response improvement, while the Judge adapts using known preference pairs derived from the Solver’s outputs rather than its own scores. Experiments show J‑Zero surpasses baselines by 4.2 points on verifiable tasks and 8.0 points on unverifiable tasks, maintaining improvement over ten iterations versus baseline degradation after two.

arXiv AI
Aug 28

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

J-Zero introduces a unified Challenger–Solver–Judge co‑evolution framework that enables self‑improvement of language models without requiring external supervision. The Challenger generates increasingly difficult tasks, the Solver learns to produce better responses, and the Judge adapts using preference pairs derived from the Solver’s own outputs rather than from its own scores. Experiments show J‑Zero outperforms baselines by an average of 4.2 points on verifiable tasks and 8.0 points on unverifiable tasks, maintaining improvement over ten iterations while baselines degrade after two.

By Gyouk Chu, Myeongho Jeon, Eunho Yang
arXiv AI
Jun 26

The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators

arXiv:2606. 26294v1 Announce Type: cross Abstract: Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains.

By Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccol\`o Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane
arXiv AI
3d ago

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

Rep2Skill introduces a representation-guided framework that enables large language model agents to self-evolve their textual skills by analyzing internal representation trajectories from agent rollouts. The method identifies execution turns that deviate from successful dynamics and uses these signals, together with execution contexts, as actionable feedback for targeted skill revision. Experiments with two open-source LLMs across two agent environments demonstrate that Rep2Skill consistently outperforms purely text-based approaches, showing that incorporating internal representations can enhance agent self-improvement.

By Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
arXiv AI
3d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu