arXiv AI By Shoupeng Wang, Jiantao Qiu, Wuyang Zhang, Conghui He

Co-Scraper: query-aware DOM Pruning and Reusable Scraper Synthesis for Lightweight Web Data Extraction

Read the original on arXiv AI →

arXiv:2606. 14821v1 Announce Type: cross Abstract: The abundant and heterogeneous nature of web content necessitates automated information extraction, and generating scrapers that can be reused across similar web pages offers an effective solution for scalable data extraction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 31

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents

The paper introduces Sieve, a search‑inspect‑fetch framework that leverages a Boolean Query Language (BQL) to target specific webpage fields, rank candidates, present structure‑rich result cards, and fetch only selected sections. Compared to traditional Search‑Visit agents, Sieve achieves higher accuracy across three QA collections while reducing token usage by 20.7–50.6%. Boolean filtering consistently improves performance for all tested rankers and remains effective across different retrievers and agent backbones.

By Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon
arXiv Machine Learning
Sep 2

Web Price Extraction: State of the Art and an Adaptive Browserless Implementation

The paper presents an adaptive browserless system for extracting prices from e‑commerce websites, improving robustness to structural differences. It combines HTML fragmentation with syntactic, semantic, and frequency rules, then enhances the baseline with a Bayesian weight‑update and a genetic‑algorithm parameter optimizer. The hybrid approach raises precision from 77.2% to 87.3% and cuts per‑page processing time by about 14%, positioning it as a competitive, low‑cost alternative to manual wrappers, browser‑based, or LLM‑driven methods.

By Evgeniia Kositsyna, Jorge Lloret-Gazo
arXiv AI
Jul 2

BaRA: BFS-and-Reflection Web Data Collection Agent

arXiv:2607. 00007v1 Announce Type: cross Abstract: Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, return incomplete multimodal outputs, or return media URLs that are not directly downloadable.

By Soojeong Lee, Joseph Lee, Yongseong Cho, Sunjae Kim, Youngwoo Moon, Kyungwoo Song