arXiv:2607. 14989v1 Announce Type: cross Abstract: Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction.
By Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, Runming He, Zhen Hao Wong, Meiyi Qiang, Hao Liang, Zhao Cao, Hao Jiang, Chong Chen, Wentao Zhang
arXiv:2609.36887v1 Announce Type: new
Abstract: Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component o...
By Bo Mao, Hang He, Linting Wang, Lizhi Lin, Maosen Zhou, Guanming Liu, Jinxiu Liu, Tianyu Huai, Chaoyun Zhang, Bingxuan Li, Kepeng Lei, Guanting Dong, Zhou Shao, Rui Zheng, Hang Yan, Jie Zhou, Chengcheng Wan, Tao Gui, Liang He, Xipeng Qiu
The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.
By Qiankai Xu
arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.
By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
ZGCM-1 is a 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. It uses a core premise that compact models can overcome capacity limits by combining deliberate internal thinking with active external tool use, supported by a 256K context and an end‑to‑end high‑efficiency training recipe that includes interleaved gated sliding‑window and full attention, a stable FP8 Muon optimizer, progressive curriculum scaling, and reformulation of interaction traces into Markov Decision Processes. The model is competitive with much larger frontier models on challenging mathematical reasoning and agentic search tasks, offers a ~4.2× efficiency improvement in pre‑training time‑to‑loss, and its weights, checkpoints, training code, data recipes, and logs are fully open‑source to support community research.
By Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren
arXiv:2606. 09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind.
By Sushant Mehta, Liudas Panavas, Edwin Chen