arXiv AI By Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

Read the original on arXiv AI →

arXiv:2608. 09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 24

Improving LLM-based Autonomous Web Agents with Filtering

The paper investigates how to improve large language model (LLM) based autonomous web agents by filtering irrelevant webpage content. The authors reproduce baseline models on the WebArena benchmark and identify failure modes caused by raw HTML input. They propose DeBERTa‑ and T5‑based retrieval models that rank HTML elements by relevance, fine‑tuned on Mind2Web data, and demonstrate that the DeBERTa model raises the LLaMA‑2‑70B agent’s success rate from 1.97% to 2.96%. Additionally, a zero‑shot ColBERT retriever achieves recall of 0.52 on Mind2Web and 0.47 on WebArena.

By Zhitong Guo, Jing Yu Koh, Ruiyu Li
arXiv AI
Jul 31

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

arXiv:2607. 26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule.

By Guanming Xiong, Penghui Zhang