The World Wide Web was built on an assumption held for three decades: the primary consumer of web content is a human being. This permeates every layer; its access model presumes human visitors, its economics rest on human attention, and its content targets human perception.
The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.
By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
arXiv:2607. 22953v1 Announce Type: new Abstract: Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into question a model where third-parties collect and control massive amounts of user data.
By Sourena Khanzadeh, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
arXiv:2607. 08147v1 Announce Type: cross Abstract: Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces.
By Corban Villa, Alp Eren Ozdarendeli, Sijun Tan, Raluca Ada Popa
An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files.
arXiv:2606. 30801v1 Announce Type: cross Abstract: Personalization algorithms determine what content users encounter on online platforms.
By Alessandro Morosini, Sarah H. Cen, Andrew Ilyas, Hedi Driss, Aleksander M\k{a}dry, Chara Podimata
arXiv:2607. 13718v1 Announce Type: cross Abstract: As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail.
By Alexandra E. Michael, Franziska Roesner
arXiv:2608. 03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself.
By Holly Lewis (Southern Illinois University Carbondale)
arXiv:2607. 05277v1 Announce Type: cross Abstract: Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data.
By Kristina Nikoli\'c, Egor Zverev, Javier Rando, Matthew Jagielski, Edoardo Debenedetti, Florian Tram\`er
arXiv:2606. 14027v1 Announce Type: cross Abstract: Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instructions.
By Xilong Wang, Xiaoxing Chen, Patrick Li, Dawn Song, Neil Gong
arXiv:2608.22061v1 Announce Type: new
Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and th...
By Minjae Seo, Wonwoo Choi, Geonwoo Han, Taekyoung Kwon, Yongsu Kim, Sang Seo, Jaewon Noh, Hankyul Baek, Seongyun Seo, Myoungsung You
arXiv:2603.17170v2 Announce Type: replace-cross
Abstract: AI agents increasingly execute users' natural-language (NL) tasks by calling Web services, yet today's Web authorizes these calls through OAu...
By Reshabh K Sharma, Linxi Jiang, Shuo Chen, Zhiqiang Lin