The paper investigates how to improve large language model (LLM) based autonomous web agents by filtering irrelevant webpage content. The authors reproduce baseline models on the WebArena benchmark and identify failure modes caused by raw HTML input. They propose DeBERTa‑ and T5‑based retrieval models that rank HTML elements by relevance, fine‑tuned on Mind2Web data, and demonstrate that the DeBERTa model raises the LLaMA‑2‑70B agent’s success rate from 1.97% to 2.96%. Additionally, a zero‑shot ColBERT retriever achieves recall of 0.52 on Mind2Web and 0.47 on WebArena.
By Zhitong Guo, Jing Yu Koh, Ruiyu Li
arXiv:2606. 06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance.
By Shiyun Xiong, Dongming Wu, Peiwen Sun, Yuang Ai, Bokang Yang, Wencheng Han, Xiao-Hui Li, Xiangyu Yue
arXiv:2603.18272v2 Announce Type: replace
Abstract: While large language models (LLMs) have advanced the development of general-purpose agents, robust generalization to unseen tasks remains challengi...
By Thomas Palmeira Ferraz, Romain Deffayet, Vassilina Nikoulina, Herv\'e D\'ejean, St\'ephane Clinchant
arXiv:2605. 02411v2 Announce Type: replace Abstract: A semantic gap separates how users describe tasks from how tools are documented.
By Kyle Zheng, Han Zhang, Renliang Sun, Chenchen Ye, Wei Wang
arXiv:2605.05726v2 Announce Type: replace
Abstract: As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical...
By Hongcheol Cho, Ryangkyung Kang, Youngeun Kim
arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.
By Guanming Xiong, Penghui Zhang
arXiv:2606. 09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite.
By Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
arXiv:2606. 19911v1 Announce Type: new Abstract: The decentralized deployment of LLM agents with diverse capabilities across diverse tasks motivates infrastructure for knowledge sharing across heterogeneous agent populations.
By To Eun Kim, Xuhong He, Dishank Jain, Ambuj Agrawal, Negar Arabzadeh, Fernando Diaz
arXiv:2602.03318v4 Announce Type: replace
Abstract: Operations Research (OR) relies on expert-driven modeling--a slow and fragile process ill-suited to novel scenarios. While large language models (L...
By Yifan Shi, Jiayi Wang, Minyi Wu, Ye Fan, Jialong Shi, Jianyong Sun
arXiv:2510. 15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke.
By Pavan C Shekar, Aswanth Krishnan
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
By Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang
Toollery is a training‑free framework that compresses candidate lists for large language model agents, enabling efficient selection from thousands of skills and tools. It generates user‑intent queries from each skill or tool specification, builds a retrieval index, and limits online selection to a compact top‑k set before the LLM makes its final decision. Evaluations on the SkillRouter benchmark, BFCL‑V4, and a proprietary smart‑cockpit dataset show that Toollery improves recall and end‑to‑end selection while keeping selection costs bounded.
By Xiangxi Tian, Ran Guan