arXiv Machine Learning

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

arXiv:2606. 09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline.

arXiv AI
Aug 25

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

The paper introduces FORGE, a benchmark that rewrites real product pages into fake ones to test how often search‑augmented large language models (LLMs) recommend these polluted items. Across 12 commercial and open‑weight LLMs, a single polluted page can lead to up to 27% of recommendations being fake, rising to 73.8% when the top‑3 replacements are used. The study finds that reasoning does not help and existing defenses—skepticism prompts, consensus filters, and credibility re‑ranking—are largely ineffective.

By Minghao Luo, Liang Chen
arXiv Machine Learning
Sep 3

Training seeds and model-selection stability in recommender-system evaluation

The paper investigates how the choice of random training seed affects recommender‑system experiments. By fixing the data split and varying seeds across hyperparameter settings, the authors analyze seed effects on user‑level metrics, validation‑based model selection, and recommendation‑list agreement. Their findings show that seed variation can be detectable and its impact depends on configuration separation, validation‑to‑test transfer, and top‑k list similarity, indicating that single‑seed results may overstate evaluation stability.

By Juan Manuel Rodriguez, Oleg Lesota, Antonela Tommasel
arXiv Machine Learning
Sep 25

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.

By Dmitrij \.Zatuchin
Hugging Face Trending Papers
Sep 17

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.

arXiv Computation and Language
Sep 11

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

RAG-Safety-Bench is a benchmark designed to evaluate how retrieval-augmented generation (RAG) affects the safety of large language models (LLMs). It isolates safety impacts by testing four conditions: non-RAG, RAG with an oracle document, RAG with related but non-answer documents, and RAG with random safe documents. Results on five open-source LLMs reveal an inverse relationship between benign and unsafe capabilities, show that baseline safety guardrails do not guarantee safety in RAG, and confirm that even benign documents can trigger unsafe generation.

By Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser
Hugging Face Trending Papers
Sep 10

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

RAG-Safety-Bench is a benchmark designed to evaluate how retrieval-augmented generation (RAG) affects the safety of large language models (LLMs). It isolates safety impacts by testing four conditions: non-RAG, RAG with an oracle document, RAG with related but non-answer documents, and RAG with random safe documents. Results on five open-source LLMs reveal an inverse relationship between benign and unsafe capabilities, show that standard safety guardrails do not guarantee safety in RAG, and confirm that even benign documents can trigger unsafe outputs.

arXiv AI
Sep 25

Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

The paper investigates prompt injection attacks on Jev, a non‑generative decision model, using 510 reconstructed cases. It finds that malicious prompts can shift Jev’s action probabilities, though rarely cause it to choose the attacker’s target. Techniques such as override markers mitigate influence, while adaptive attacks that use score feedback roughly double the highest attacker‑target probability and increase success rates on new validation calls from 1.8% to 3.5%.

By Tiantong Wu, Wei Yang Bryan Lim