arXiv:2609.05576v1 Announce Type: new
Abstract: The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across statefu...
By Yirong Zeng, Shen You, Jinhang Feng, Yufei Liu, Xiao Ding, Yutai Hou, Hao Cong, Yuxian Wang, Wu Ning, Wang Xu, Bibo Cai
arXiv:2606. 01279v1 Announce Type: new Abstract: AI agents are increasingly being tasked with automating AI research itself, particularly the critical post-training phase that transforms base LLMs into aligned assistants.
By Zhengyang Zhao, Shengjie Ye, Lu Ma, Hao Liang, Hengyi Feng, Wentao Zhang
arXiv:2606. 03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data.
By Jinnuo Liu, Yue Peng, Jinhan Niu, Hongyi Wen
The paper investigates the use of looped language models for compositional tool calling, where models must coordinate multiple API calls and maintain state across interactions. Experiments on API-Bank, BFCL, and NESTful show that recurrent computation generally improves compositional and dependency-aware tool use, with accuracy increasing as recurrent depth grows. Adaptive inference offers a better compute‑performance trade‑off by allocating extra computation only when necessary.
By Andrei Cristian Popescu, Haitz S\'aez de Oc\'ariz Borde, Pietro Li\`o
arXiv:2607. 01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society.
By Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han, Guoliang Li, Ying Yan, Peng Zhang, Yu Su, Xiang Qi, Baolin Sun, Chengyuan Yang, Tao Fang, Huaiyu Ruan
arXiv:2607. 23722v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hidden information, composing tool calls, and committing state changes.
By Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
arXiv:2604. 18543v4 Announce Type: replace Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale.
By Xirui Li, Ming Li, Ion Stoica, Cho-Jui Hsieh, Tianyi Zhou
The paper introduces KOPA-Bench, a benchmark of 145 real-world tasks that evaluate multi-step tool‑calling over Korean open public APIs. It also presents EDGE, a data‑synthesis method that builds an execution‑grounded dynamic graph to generate executable multi‑step trajectories, and shows that a fine‑tuned 9B model performs nearly as well as an untuned 27B model on KOPA‑Bench and the BFCL benchmark.
By Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim
arXiv:2604. 13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments.
By Xiang Long, Li Du, Yilong Xu, RongJian Xu, Qiyanhui Lu, Ying Gao, Qinhua Xie, Fangcheng Liu, Ning Ding, Haoqing Wang, Ziheng Li, Changjiang Zhou, Jianyuan Guo, Yehui Tang
SimCRAFT is a model‑agnostic framework that distills remote sensing orchestration into a compact 7B‑scale model. It creates a large, constraint‑validated workflow planning corpus (SimRS‑14k) using a multi‑agent synthesis engine and a Mock Execution Engine, then fine‑tunes the model with Contextual Retrieval‑Augmented Fine‑Tuning (CRAFT) to reason analogically. Experiments show SimCRAFT‑7B outperforms open‑weight LLMs and rivals advanced closed‑source models, providing a lightweight, efficient baseline for autonomous remote sensing deployment.
By Haoran Wang, Jing Yao, Xu Yang, Zeqing Wang, Yang Zhang, Pedram Ghamisi, Zhengchao Chen
EDGEGEN is a synthetic task generation framework that extracts compliance rules from a tool‑calling agent’s specification to create database‑grounded edge‑case tasks that violate those rules. By combining EdgeGen with existing synthetic data generation methods, it forms a fully automated closed‑loop system that requires no human annotation. Experiments show that finetuning on EdgeGen data improves performance by 2–42 % on the tau2bench airline domain, while harness optimization yields 10–30 % gains over human‑curated and base harnesses for the Gemma‑4‑e4b model.
By Harshavardhan Abichandani, Penny Chong, Jiyuan Shen, Gunraj Singh, Ashutosh Hathidara, Marcus Duigan Xing Yu, Jane Lo, Atin Ghosh, Yipeng Li, Daniel Dahlmeier
arXiv:2605. 30407v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.
By Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng