arXiv:2606. 13610v1 Announce Type: cross Abstract: Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content.
By Minghao Luo, Liang Chen
The paper introduces FORGE, a benchmark that rewrites real product pages into fake ones to test how often search‑augmented large language models (LLMs) recommend these polluted items. Across 12 commercial and open‑weight LLMs, a single polluted page can lead to up to 27% of recommendations being fake, rising to 73.8% when the top‑3 replacements are used. The study finds that reasoning does not help and existing defenses—skepticism prompts, consensus filters, and credibility re‑ranking—are largely ineffective.
By Minghao Luo, Liang Chen
arXiv:2606. 17443v1 Announce Type: new Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel.
By Xi Chu, Yupeng Hou
arXiv:2606. 28356v1 Announce Type: cross Abstract: Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems.
By Qianfeng Wen, Yifan Simon Liu, Xin Liu, Difan Jiao, Blair Yang, Junda Wu, Zhenwei Tang
The paper investigates how the choice of random training seed affects recommender‑system experiments. By fixing the data split and varying seeds across hyperparameter settings, the authors analyze seed effects on user‑level metrics, validation‑based model selection, and recommendation‑list agreement. Their findings show that seed variation can be detectable and its impact depends on configuration separation, validation‑to‑test transfer, and top‑k list similarity, indicating that single‑seed results may overstate evaluation stability.
By Juan Manuel Rodriguez, Oleg Lesota, Antonela Tommasel
The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.
By Dmitrij \.Zatuchin