arXiv:2609.22308v1 Announce Type: new
Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing researc...
By Boyu Qiao, Zixin Tang, Xiaoshuai Hao, Wenbo Li
arXiv:2607. 05465v1 Announce Type: cross Abstract: Complex image creation and editing often require more than a single generation or editing model.
By Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu, Xiao Yao, Xiaoyang Ye, Lin Ma, Wenhao Jiang
arXiv:2609.36593v1 Announce Type: cross
Abstract: Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be des...
By Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li
VideoWeaver is an agent harness and benchmark for long‑video generation that evaluates and evolves reusable skills. It allows an agent to dynamically compose foundation skills into a workflow based on a single high‑level instruction, rather than following a static pipeline. The benchmark includes 16 task categories and 285 cases with multimodal references, and an evidence‑grounded agent‑as‑judge inspects execution traces and final videos to guide skill evolution, improving both process and output quality.
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.
By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
arXiv:2606. 14397v1 Announce Type: new Abstract: As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities.
By Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh Padarha, Kumail Alhamoud, Zihao Fu, William Lugoloobi, Kai Rawal, Hanna Yershova, Xander Davies, Taras Rumezhak, Guohao Li, Fazl Barez, Baoyuan Wu, Arkadiusz Drohomirecki, Yarin Gal, Chris Russell, Christopher Summerfield, Adam Mahdi, Volodymyr Karpiv, Philip Torr, Adel Bibi