The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.
By Dmitrij \.Zatuchin
The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.
By Dmitrij \.Zatuchin
arXiv:2608.30023v1 Announce Type: cross
Abstract: Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying...
By Dmitrij \.Zatuchin, Daniil Dzemesjuk
arXiv:2603. 08924v2 Announce Type: replace-cross Abstract: AI-powered answer engines are inherently non-deterministic: identical queries submitted at different times can produce different responses and cite different sources.
By Ronald Sielinski
arXiv:2609.34951v2 Announce Type: replace
Abstract: In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training...
By Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef
arXiv:2608.30052v1 Announce Type: cross
Abstract: When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the prod...
By Dmitrij \.Zatuchin
ZooWork-ShopRanker is a family of open e‑commerce rerankers (0.6B, 4B, and 8B) that align with human shopping preferences by using large language models as preference oracles to generate training pairs. The flagship 8B model serves as a teacher for the smaller 4B and 0.6B models, which are further refined on judged pairs. A new benchmark, ShopRank‑Bench, contains ~10,000 private‑traffic preference pairs and shows that all ZooWork models outperform the strongest open reranker baseline and their own un‑aligned versions.
By Siqiao Xue, Shuxuan Liu, Ning Hu
arXiv:2605. 12887v2 Announce Type: replace-cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes.
By Hengwei Ye, Jiasheng Mao, Zhenhan Guan, Zheng Tian
The paper evaluates how modern large language models use internal web search to answer factual questions. Using 783 static queries and 288 dynamic queries, the authors find that enabling retrieval improves accuracy on static questions but hurts confidence calibration. On dynamic queries, models often retrieve but still achieve less than 70% accuracy, mainly due to poor query formulation and source selection, indicating that internal web search works better as a quick verification tool than a full information‑retrieval system.
By Sahil Kale
TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.
By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan
Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.
arXiv:2609.05766v1 Announce Type: cross
Abstract: The dominant approach to building domain-specific pretraining corpora is to filter large web archives such as CommonCrawl. This works well for popula...
By Chirag Garg, Eelaaf Zahid, Farhan Ahmed, Jay Pankaj Gala, Eric Butler, Heiko Ludwig