The paper introduces FORGE, a benchmark that rewrites real product pages into fake ones to test how often search‑augmented large language models (LLMs) recommend these polluted items. Across 12 commercial and open‑weight LLMs, a single polluted page can lead to up to 27% of recommendations being fake, rising to 73.8% when the top‑3 replacements are used. The study finds that reasoning does not help and existing defenses—skepticism prompts, consensus filters, and credibility re‑ranking—are largely ineffective.
By Minghao Luo, Liang Chen
arXiv:2609.06027v1 Announce Type: cross
Abstract: Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. E...
By Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong
arXiv:2606. 28356v1 Announce Type: cross Abstract: Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems.
By Qianfeng Wen, Yifan Simon Liu, Xin Liu, Difan Jiao, Blair Yang, Junda Wu, Zhenwei Tang
arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.
By Hyunseok Paeng
arXiv:2606. 17443v1 Announce Type: new Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel.
By Xi Chu, Yupeng Hou
The paper demonstrates that large language models (LLMs) used for forecasting real‑world events can be manipulated by simply publishing new articles, even without direct access to the model or its retriever. By injecting a small number of targeted news pieces into a common crawl corpus, an adversary can flip over half of the forecast probabilities and significantly degrade forecast accuracy. The study also shows that common defense strategies can be cheaply bypassed, highlighting the vulnerability of probabilistic LLM judgments to information‑supply‑chain attacks.
By Yuan Lu, Yukuan Zhang
arXiv:2608. 11390v1 Announce Type: new Abstract: Generative engines are reshaping the web ecosystem by making citations a key mechanism for allocating attention, attribution, and downstream value.
By Chen Xu, Zitian Guo, Chenyan Xiong
arXiv:2603. 00801v2 Announce Type: replace Abstract: Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources.
By Shrey Shah, Levent Ozgur
The paper introduces GEO Defender, a two‑stage defense system designed to protect generative search engines from malicious Generative Engine Optimization (GEO) attacks that rewrite web documents to manipulate generated answers. GEO Defender comprises a Shield Reranker, which learns a defensive residual to demote GEO‑rewritten documents while maintaining relevance, and a Training‑Free Shield Generation component that creates a natural‑language library guiding the target LLM’s source usage during inference. Experiments on both closed‑source and open‑source large language models show that GEO Defender dramatically lowers attack success rates from 50.32% to 6.20%, preserves over 94% of benign evidence usage, and maintains answer quality while generalizing to unseen attacks.
By Haozhang Li, Yangguang Shao, Xinjie Lin, Zhong Guan, Mi Zhou, Junzheng Shi
arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.
By Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo
The paper investigates why misinformation spreads more quickly on engagement‑based platforms by dissecting the recommendation algorithm of X. It identifies an engagement fungibility mechanism that rewards instant reactions (likes, retweets) over thoughtful engagement (replies, quotes), allowing misinformation—which tends to attract instant reactions—to receive more recommendations. The authors validate this mechanism through a simulation on the USC X 2024 election corpus, showing that adjusting metric weights has little effect, while requiring thoughtful engagement before amplification can significantly reduce the credibility exposure gap without harming mainstream content or engagement.
By Pan Li, Shuang Gao
arXiv:2609.38270v1 Announce Type: cross
Abstract: Advancing beyond traditional static scoring models, LLM-powered agentic recommender systems (LLM-ARS) instantiate users and items as autonomous agent...
By Yurong Hao, Wen Zhou, Guowei Guan, Tiantong Wu, Fuyao Zhang, Wei Yang Bryan Lim