arXiv:2510. 11560v2 Announce Type: replace-cross Abstract: The advent of LLMs has given rise to generative search, a new search paradigm in which LLMs retrieve information from the web related to a query and synthesize it into a single, coherent response.
By Elisabeth Kirsten, Jost Grosse Perdekamp, Qinyuan Wu, Mihir Upadhyay, Krishna P. Gummadi, Muhammad Bilal Zafar
Agent2UCB is a new agentic system designed for Generative Engine Optimization (GEO), which refines content to boost its likelihood of being cited or summarized by generative AI search engines. The system autonomously evaluates nine GEO strategies for each content item, selects the most effective one, and speeds up this selection using a bandit-based Agent2UCB policy that blends large language model priors with real-time reward signals. Additionally, it offers a lightweight, text-only SEO readiness check that assesses readability, topical coverage, and EEAT-style credibility, and experiments on GEO-Bench demonstrate consistent visibility gains while maintaining SEO quality.
By Sheldon Yu, Rui Wang, Tong Yu, Sungchul Kim, Doga Dogan, Junda Wu, Julian McAuley
arXiv:2602. 12187v2 Announce Type: replace-cross Abstract: Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generative capabilities to deliver synthesized answers.
By Sunghwan Kim, Wooseok Jeong, Serin Kim, Sangam Lee, Dongha Lee
arXiv:2608. 16824v1 Announce Type: new Abstract: Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines.
By Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen, Yang Zhang
arXiv:2606. 02814v1 Announce Type: cross Abstract: Neural retrievers are trained to estimate query-document relevance from annotated query-document pairs.
By Francisco Valentini, Edgar Altszyler, Martin Fajcik
Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.
By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen