arXiv:2607. 20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time.
By Bronislav Sidik, Chaya Levi, Nizzan Kimhi
arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.
By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv:2607. 14890v1 Announce Type: new Abstract: Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-merge remain claims unless supported by current evidence.
By Jek Huang, Jeffery Hsia, Jiayi Sun, Freddie Shi, Wei Huang, Ian H. White
arXiv:2606. 31763v1 Announce Type: new Abstract: Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experimental feedback must remain aligned from protocol and SOP design to code and physical execution.
By Yankai Jiang, Weiting Tang, Haoran Sun, Zhenyu Tang, Yuejie Hou, Yingnan Han, Rubo Wang, Yueyuxiao Yang, Cheng Liang, Lilong Wang, Wenjie Lou, Xiaosong Wang, Lei Bai, Meng Yang
arXiv:2609.24165v1 Announce Type: new
Abstract: Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound...
By Pawan K. Tripathi, Hemant Sharma, Andrew Chuang, Mathew J. Cherukara
PentestChain is a ten‑phase automated penetration testing framework that uses a cost‑aware AI cascade, starting with a local 7B‑parameter Ollama model (qwen2.5‑7b) and then free‑tier OpenRouter and Cerebras models, with a rule‑based fallback. It exposes the entire pipeline through a Model Context Protocol (MCP) server that includes eleven tools. The authors evaluate the framework using standard testbeds (AutoPenBench, Cybench subset, PentestGPT 182‑sub‑task benchmark) and report that the local model keeps paid‑API cost at zero while detecting 26 services and enriching 34 CVEs on legacy targets.
By Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem