arXiv Machine Learning

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.

arXiv Computation and Language
Sep 7

Repeated Queries Exhaust an LLM's Brand Recommendations but Not Its Sources

The study examines how repeated identical buying questions affect the brand recommendations of large language models (LLMs) with and without web‑search retrieval. Across 300 question‑engine cells, five engines that did not use web search continued to add new, previously unseen brands up to run 15, while the single retrieval‑enabled engine’s list plateaued earlier. Domain citations continued to grow throughout the runs, indicating that LLMs keep accumulating source diversity even as brand lists stabilize.

By Dmitrij \.Zatuchin
arXiv Computation and Language
Sep 17

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

arXiv:2609.18729v1 Announce Type: cross Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...

By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
arXiv Machine Learning
Jun 9

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.

By Hyunseok Paeng
arXiv Computation and Language
Sep 25

Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit

The study audits large language model (LLM) outputs by measuring how well repeated queries recover a collected set of responses versus the full set of possible outputs. Using sample-based rarefaction on 4,500 responses from 50 buying questions across six configurations, the authors find historical-dictionary median recovery rates between 92.6% and 95.2%, which drop to 89.5%–94.7% after re‑adjudicating all candidate strings. Additional analyses with Gemini 3.1 Pro annotations and matched roster data confirm that recovery percentages vary with extraction methods, question selection, and the finite reference collection, underscoring the need for explicit measurement definitions and sensitivity analyses in LLM audits.

By Dmitrij \.Zatuchin
arXiv Computation and Language
Sep 4

The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations

The Dice Roll Method is a standardized protocol for auditing large language model brand recommendations through repeated queries. It decomposes total response variance into sampling, prompt‑phrasing, run‑to‑run, and model‑version components, and uses a negative‑binomial mixed model, Cliff’s delta, and bootstrap techniques to guide iteration counts. The study identifies three iteration tiers—exploratory (n=5), confirmatory (n=10), and rigorous (n=15)—and recommends a compact battery of four complementary metrics for robust evaluation.

By Dmitrij \.Zatuchin