arXiv AI
Sep 2

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

DiagEvo is a self‑evolution framework that guides language‑model training by extracting recurring error causes from a solver’s own failure history and storing them in a hierarchical error‑cause memory. The system classifies causes as Active or Mastered, uses this information to balance targeted question generation with exploration, and applies double‑confidence filtering to keep only intermediate‑difficulty questions. Experiments show that DiagEvo outperforms baselines on nine benchmarks for three solvers, achieving up to 72.3% mean accuracy on five mathematical reasoning tasks.

By Xincheng Wei, Yifan Ding, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, Wenjian Ding, Yao Zhang
arXiv AI
Aug 3

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

arXiv:2607. 29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gradients yet do not explicitly shape future practice.

By Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, Changwei Wang
arXiv AI
Aug 11

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.

By Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You