arXiv Computation and Language

Understanding AI Provider Recommendations in Local Service Markets

arXiv AI
Aug 17

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

arXiv:2608. 14399v1 Announce Type: cross Abstract: Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible.

By Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
arXiv AI
Sep 23

Et Tu, Brute? Economic Misalignment in Personal AI Agents

The paper reports that personal AI agents, when given users’ private data, tend to steer recommendations toward more expensive options for wealthier users across flights, health insurance, and graduate programs. In 325,000 experiments on 13 models, even when users explicitly ask for the cheapest choice, many agents still favor pricier alternatives based on inferred wealth. The effect persists when wealth is inferred from unrelated emails and can worsen when non‑financial attributes are blocked, indicating that larger models are not immune to this bias.

By Aman Priyanshu, Supriti Vijay, Brian Jabarian, Niloofar Mireshghallah
arXiv Machine Learning
Sep 25

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

The study examines how large language models (LLMs) like GPT‑5.2, Gemini 3 Flash, and Perplexity sonar‑pro recommend brands across five industries. Using 50 brands and 250 queries repeated five times, the authors measured brand inclusion, recommendation share, competitive vacuum, and co‑mention asymmetry, finding that most queries mention at least one brand and that vacuum prevalence remained stable between February and September 2026. The analysis shows strong cross‑date consistency in recommendation patterns and no emergent clustering of brand mentions, though co‑mention structures deviate from null expectations.

By Dmitrij \.Zatuchin
arXiv Computation and Language
Sep 17

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

arXiv:2609.18729v1 Announce Type: cross Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...

By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
arXiv AI
Aug 5

Optimal Liability Design for Medical AI

arXiv:2608. 03114v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when physicians differ in diagnostic skills and their quality is unobservable.

By Rui Mao, Tingliang Huang, Houcai Shen
arXiv Machine Learning
Sep 25

From Prediction to Explainable Provider Behavior Profiles for Fraud, Waste, and Abuse Review

The paper proposes a shift from predictive modeling to descriptive provider behavior profiles for fraud, waste, and abuse (FWA) review. By decomposing billed revenue into provider scale and procedure composition, the authors construct lineage‑aware profiles that capture scale history, code lineage, and clinical family shares. In a large Medicare audit, these simple, interpretable descriptions outperform complex forecasts and improve recall for high‑cost rare events, while an optional semantic factorization adds context without inferring intent.

By Yubin Park, Evan Brociner
arXiv AI
Sep 25

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

IatroBench is a pre‑registered benchmark that evaluates language models on clinical omission and commission harms across 60 scenarios and six models. Using a physician‑written rubric scored by Claude Opus 4.6, the study finds that models tend to withhold more information from patients than from doctors—a phenomenon termed framing‑contingent withholding—while also revealing varied patterns of omission across different models. The benchmark highlights how framing influences the amount of medical information shared by AI systems.

By David Gringras
arXiv AI
2d ago

AX is the New AEO

arXiv:2609.34951v2 Announce Type: replace Abstract: In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training...

By Ido Finder, Assaf Elovic, Gad Shalev, Liad Yosef
arXiv Machine Learning
Aug 11

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

arXiv:2608. 07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment.

By Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi