The study examines how repeated identical buying questions affect the brand recommendations of large language models (LLMs) with and without web‑search retrieval. Across 300 question‑engine cells, five engines that did not use web search continued to add new, previously unseen brands up to run 15, while the single retrieval‑enabled engine’s list plateaued earlier. Domain citations continued to grow throughout the runs, indicating that LLMs keep accumulating source diversity even as brand lists stabilize.
By Dmitrij \.Zatuchin
arXiv:2606. 17443v1 Announce Type: new Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel.
By Xi Chu, Yupeng Hou
arXiv:2609.18729v1 Announce Type: cross
Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...
By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
arXiv:2608.30023v1 Announce Type: cross
Abstract: Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying...
By Dmitrij \.Zatuchin, Daniil Dzemesjuk
arXiv:2606. 26116v1 Announce Type: cross Abstract: A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per provider?
By Will Jack, Noah Lehman, Keller Maloney, Sarah Xu
arXiv:2609.18341v1 Announce Type: cross
Abstract: When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recom...
By Hazem Ibrahim, Yasir Zaki
arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.
By Hyunseok Paeng
The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.
By Dmitrij \.Zatuchin
The Dice Roll Method is a standardized protocol for auditing large language model brand recommendations through repeated queries. It decomposes total response variance into sampling, prompt‑phrasing, run‑to‑run, and model‑version components, and uses a negative‑binomial mixed model, Cliff’s delta, and bootstrap techniques to guide iteration counts. The study identifies three iteration tiers—exploratory (n=5), confirmatory (n=10), and rigorous (n=15)—and recommends a compact battery of four complementary metrics for robust evaluation.
By Dmitrij \.Zatuchin
arXiv:2608. 10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog.
By Srijith Ravikumar
arXiv:2608.30052v1 Announce Type: cross
Abstract: When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the prod...
By Dmitrij \.Zatuchin
arXiv:2603. 08924v2 Announce Type: replace-cross Abstract: AI-powered answer engines are inherently non-deterministic: identical queries submitted at different times can produce different responses and cite different sources.
By Ronald Sielinski