Simon Willison
6d ago

Quoting Muse AI Agent

The article reports a failed MX Keys Mini pickup where a customer, Usman, arrived at the building at 9:15 but was not met, leading to a negative rating. The author acknowledges that an auto‑reply incorrectly confirmed the author's presence at 9:27, worsening the situation, and has apologized on behalf of the account. They are considering disabling auto‑replies that promise the author is home when they cannot confirm it.

arXiv AI
Sep 4

Identifying AI Web Scrapers Using Canary Tokens

The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.

By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
Simon Willison
Sep 7

Creepy crawlies

Konstantin Ryabitsev highlights the growing problem of abusive web crawlers that consume excessive CPU resources on git.kernel.org, the official Git repository for the Linux kernel. He notes that at any given moment, 14 CPU cores across five geo‑distributed nodes are dedicated solely to rendering git commits as HTML for these scrapers, surpassing the CPU usage for all legitimate access such as git clones. This issue raises concerns for services like Datasette, which also serve large numbers of crawlable web pages.