arXiv:2603. 12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings.
By Guanyu Jiang, Zhaochen Su, Xiaoye Qu, Yi R. Fung
arXiv:2608.30760v1 Announce Type: new
Abstract: Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual obse...
By Ziyi Bai, Siqi Li, Tinglei Huang, B\"orje F. Karlsson
arXiv:2605. 13527v3 Announce Type: replace Abstract: Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines.
By Kangning Zhang, Shuai Shao, Qingyao Li, Jianghao Lin, Lingyue Fu, Shijian Wang, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
arXiv:2609.37810v1 Announce Type: cross
Abstract: Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains chal...
By Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang
arXiv:2609.40048v1 Announce Type: new
Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget,...
By Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress. Raw interaction traces preserve such information but are long and noisy to condition on, whereas text-only skills often omit the visual state that makes a procedure applicable.
arXiv:2607. 01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently.
By Zongxia Li, Dawei Liu, Fuxiao Liu, Yuhang Zhou, Xiyang Wu, Jingxi Chen, Jing Xie, Xiaomin Wu, Lichao Sun
arXiv:2609.15683v1 Announce Type: new
Abstract: While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particu...
By Ziqian Fan, Shibo Xu, Junjie Li, Xiangyu Zhao, Shengyuan Ding, Yifan Yang, Zhenjie Yang, Haodong Duan, Yue Zhou, Zhihang Zhong, Xue Yang
SkillRL is a framework that enhances large language model agents by automatically discovering and evolving skills from raw experience. It builds a hierarchical skill library called SkillBank, uses an adaptive retrieval strategy for heuristics, and allows the skill library to co‑evolve with the agent’s policy during reinforcement learning. These techniques reduce token usage and improve reasoning, achieving state‑of‑the‑art results on ALFWorld, WebShop, and seven search‑augmented tasks, outperforming baselines by 15.3% and remaining robust as task complexity grows.
By Peng Xia, Jianwen Chen, Hanyang Wang, Jiaqi Liu, Kaide Zeng, Yu Wang, Siwei Han, Yiyang Zhou, Xujiang Zhao, Haifeng Chen, Zeyu Zheng, Cihang Xie, Huaxiu Yao
arXiv:2608. 10775v1 Announce Type: new Abstract: Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress.
By Zhou Liu, Ligang Huang, Zeli Su, Zewei Pan, Zhaoyang Han, Xing Chen, Yuanfeng Song, Wentao Zhang
arXiv:2608. 05970v1 Announce Type: cross Abstract: Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks.
By Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tang, Jiwen Lu
arXiv:2610.10782v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs),
but it typicall...
By Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang