arXiv AI By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger

Identifying AI Web Scrapers Using Canary Tokens

Read the original on arXiv AI →

The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

arXiv:2609.06027v1 Announce Type: cross Abstract: Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. E...

By Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong
arXiv Computation and Language
Sep 1

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

arXiv:2510.03204v2 Announce Type: replace Abstract: Web agents powered by large language models (LLMs) must process lengthy web page observations to complete user goals; these pages often exceed tens...

By Imene Kerboua, Sahar Omidi Shayegan, Megh Thakkar, Xing Han L\`u, L\'eo Boisvert, Massimo Caccia, J\'er\'emy Espinas, Alexandre Aussem, V\'eronique Eglin, Alexandre Lacoste