arXiv AI

Cheap, open agents make LLM pollution harder to mitigate

The paper titled "Cheap, open agents make LLM pollution harder to mitigate" reports that open-weight language models combined with open-source agentic frameworks can produce synthetic survey responses that are competitive with commercial agents and harder to detect. The authors compared nine agent configurations, finding that fully open agents run locally without usage fees and that no single detection check reliably identifies all agents. Open-text responses were the most effective at distinguishing agents from humans, highlighting the need for multilayered detection strategies.

arXiv Computation and Language
Sep 16

Towards Detecting AI-Assisted Responses in Online Surveys

The paper introduces ASURRE, a benchmark dataset for detecting AI‑assisted responses in online surveys. It evaluates how different LLM usage strategies—ranging from full generation to persona‑grounded agentic completion—affect the performance of existing machine‑generated text detectors. The study finds that while naive AI usage is easily detected, more sophisticated persona‑grounded agents approach chance performance, yet still leave identifiable behavioural traces that can be aggregated to improve detection.

By Qizhou Wang, Bogdan Mamaev, Christopher Leckie
arXiv AI
Sep 10

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

arXiv:2609.06027v1 Announce Type: cross Abstract: Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. E...

By Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong
arXiv AI
Jun 12

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

arXiv:2606. 12737v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted external sources.

By Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox, George Lee, Jiliang Tang, Tomas Pfister, Long T. Le