arXiv AI

Co-Scraper: query-aware DOM Pruning and Reusable Scraper Synthesis for Lightweight Web Data Extraction

arXiv:2606. 14821v1 Announce Type: cross Abstract: The abundant and heterogeneous nature of web content necessitates automated information extraction, and generating scrapers that can be reused across similar web pages offers an effective solution for scalable data extraction.

arXiv Computation and Language
Aug 31

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents

The paper introduces Sieve, a search‑inspect‑fetch framework that leverages a Boolean Query Language (BQL) to target specific webpage fields, rank candidates, present structure‑rich result cards, and fetch only selected sections. Compared to traditional Search‑Visit agents, Sieve achieves higher accuracy across three QA collections while reducing token usage by 20.7–50.6%. Boolean filtering consistently improves performance for all tested rankers and remains effective across different retrievers and agent backbones.

By Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
arXiv Machine Learning
Sep 2

Web Price Extraction: State of the Art and an Adaptive Browserless Implementation

The paper presents an adaptive browserless system for extracting prices from e‑commerce websites, improving robustness to structural differences. It combines HTML fragmentation with syntactic, semantic, and frequency rules, then enhances the baseline with a Bayesian weight‑update and a genetic‑algorithm parameter optimizer. The hybrid approach raises precision from 77.2% to 87.3% and cuts per‑page processing time by about 14%, positioning it as a competitive, low‑cost alternative to manual wrappers, browser‑based, or LLM‑driven methods.

By Evgeniia Kositsyna, Jorge Lloret-Gazo
arXiv AI
Jul 2

BaRA: BFS-and-Reflection Web Data Collection Agent

arXiv:2607. 00007v1 Announce Type: cross Abstract: Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, return incomplete multimodal outputs, or return media URLs that are not directly downloadable.

By Soojeong Lee, Joseph Lee, Yongseong Cho, Sunjae Kim, Youngwoo Moon, Kyungwoo Song
arXiv Computation and Language
Aug 31

PACE: Publisher-Adaptive Content Extraction via Agentic Automation

PACE is an agentic framework that learns publisher‑specific extraction configurations from representative web pages and user requirements. During training it employs LLMs to analyze page structure and gather reusable extraction patterns, then creates a deterministic extractor template for inference that eliminates the need for further LLM calls. Experiments on article bodies, metadata, images, and tables show that PACE surpasses scalable non‑manual baselines and approaches the quality of manually engineered publisher‑specific parsers.

By Zhanlin Liu, Munirathnam Srikanth
arXiv AI
Aug 11

TreeHop: Efficient Embedding-Level Query Rewriter

arXiv:2504. 20114v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries require synthesizing information across multiple document chunks.

By Zhonghao Li, Kunpeng Zhang, Jinghuai Ou, Shuliang Liu, Xuming Hu
arXiv AI
Jun 30

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

arXiv:2606. 28344v1 Announce Type: cross Abstract: Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting.

By Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min
arXiv AI
Sep 25

WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks

WebArxiv is a reproducible benchmark designed to evaluate multimodal web agents on arXiv-related tasks. It consists of 510 static, time‑invariant tasks that require multi‑constraint paper retrieval, fine‑grained content extraction, and cross‑paper comparison, each with a deterministic ground truth. The benchmark highlights challenges for foundation‑model agents, such as over‑reliance on fixed interaction histories, and introduces a lightweight dynamic‑memory mechanism to improve adaptive retrieval and reasoning.

By Zihao Sun, Zijing Shi, Ling Chen