The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.
By Dmitrij \.Zatuchin
arXiv:2608. 10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog.
By Srijith Ravikumar
arXiv:2606. 17443v1 Announce Type: new Abstract: Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel.
By Xi Chu, Yupeng Hou
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany, $τ^2$-bench, and AppWorld), that the agent main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7-23%.
The paper demonstrates that a preference‑optimization objective can learn to distinguish reliable from unreliable sources by installing a prior‑dependent reliability switch. By training on data where a source’s stated reliability is paired with its answer, the model learns to flip its response only when the stated reliability exceeds a threshold that grows with the model’s prior. Experiments on Qwen2.5‑7B‑Instruct and Llama‑3.1‑8B show that this switch generalizes to unseen reliability values and follows stated reliability over role prestige, whereas supervised imitation fails to learn it.
By Sen Yang, Yuen-Hei Yeung
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness.
The Dice Roll Method is a standardized protocol for auditing large language model brand recommendations through repeated queries. It decomposes total response variance into sampling, prompt‑phrasing, run‑to‑run, and model‑version components, and uses a negative‑binomial mixed model, Cliff’s delta, and bootstrap techniques to guide iteration counts. The study identifies three iteration tiers—exploratory (n=5), confirmatory (n=10), and rigorous (n=15)—and recommends a compact battery of four complementary metrics for robust evaluation.
By Dmitrij \.Zatuchin
The paper introduces a new evaluation method called "same-input rerun" to assess the consistency of clinical language‑model agents across repeated runs. By replaying 1,000 MedAgentBench tasks with identical inputs, the authors find that action‑level outputs—such as test orders, medication requests, and referrals—vary significantly, even when benchmark scores remain unchanged. The study demonstrates that current benchmarks, which typically evaluate only a single run per task, can miss substantial behavioral divergence.
By Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han, Wenbin Zhang
arXiv:2608. 11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.
By Vasundra Srinivasan
The study evaluates whether aggregated brand recommendation profiles can identify the language model that generated them. Using 6,475 responses from five deployed endpoints, a character‑n‑gram classifier accurately attributes single responses to the correct system (97.84% accuracy). However, when responses are aggregated into domain‑condition units, the classifier’s performance drops to 66.53%, and a forest model misclassifies all gift‑domain units, indicating that aggregated brand behaviour does not reliably reveal the underlying system.
By Dmitrij \.Zatuchin
arXiv:2606. 27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy.
By Josef Chen
arXiv:2608. 04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why.
By Atul Anand, Sourav Chattaraj