Bluesky reply bot checker
Read the original on Simon Willison →The Flow has not summarised this story yet — read it at Simon Willison.
The Flow has not summarised this story yet — read it at Simon Willison.
The article reports a failed MX Keys Mini pickup where a customer, Usman, arrived at the building at 9:15 but was not met, leading to a negative rating. The author acknowledges that an auto‑reply incorrectly confirmed the author's presence at 9:27, worsening the situation, and has apologized on behalf of the account. They are considering disabling auto‑replies that promise the author is home when they cannot confirm it.
The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.
arXiv:2608.22061v1 Announce Type: new Abstract: Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and th...
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it.
arXiv:2609.27155v1 Announce Type: cross Abstract: With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader...
Konstantin Ryabitsev highlights the growing problem of abusive web crawlers that consume excessive CPU resources on git.kernel.org, the official Git repository for the Linux kernel. He notes that at any given moment, 14 CPU cores across five geo‑distributed nodes are dedicated solely to rendering git commits as HTML for these scrapers, surpassing the CPU usage for all legitimate access such as git clones. This issue raises concerns for services like Datasette, which also serve large numbers of crawlable web pages.