arXiv AI

Nudging Sustainable Choices through LLM-Generated Recommendation Explanations

arXiv:2607. 25726v1 Announce Type: new Abstract: Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices.

arXiv AI
Sep 3

The Utility of LLMs in Recommender Systems Explanation Evaluation

The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.

By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv AI
Sep 11

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

XAI-Arena proposes using large language models (LLMs) as judges to evaluate the quality of explainable AI (XAI) explanations, aiming for reproducibility, scalability, and multidimensional assessment. The framework assesses dimensions such as simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability across different datasets, models, and stakeholder personas. Human validation shows a strong positive correlation between LLM-generated and human ratings (Spearman's rho = .693, p < .001), supporting the viability of LLM-based evaluations.

By Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel
arXiv Computation and Language
Sep 7

From Plausible to Actionable: A Position on LLM Self-Explanations

The paper discusses how Large Language Models can produce natural language self‑explanations that appear plausible but may not accurately reflect the model’s reasoning. It critiques current evaluation methods for such explanations and offers practical guidelines to assess their plausibility and faithfulness. Additionally, it argues that evaluation should also consider the actionability of these explanations, showing how they can aid decision‑making for various stakeholders.

By Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti
arXiv AI
Sep 17

Scaling Articulated Rationales for MLLM-based Recommendation

The paper introduces SARA, an industrial framework that scales articulated user rationales (AURs) for recommendation systems. It curates a high‑quality AUR dataset from 240 M users, trains a 7B‑parameter MLLM (SARA‑7B) to generate rationales for millions of authors, and integrates these generated rationales into a production ranking model (SARA‑Ranker). Offline and online experiments demonstrate that the system produces more specific, polarity‑consistent rationales and improves user engagement while reducing negative feedback.

By Haoke Xiao, Yueyang Liu, Yuhui Zhang, Xiang Chen, Yufei Liu, Jia Xu, Yalong Guan, Xiaolan Zhu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zijie Meng, Zejian Zhang, Ruochen Yang, Xiangyu Wu, Tingting Gao, Han Li, Lantao Hu, Cheng Luo, Kun Gai
arXiv Computation and Language
4d ago

Evaluating Alignment of Behavioral Dispositions in LLMs

arXiv:2602.11328v2 Announce Type: replace Abstract: As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential. In this work, we focus on behavioral dis...

By Amir Taubenfeld, Zorik Gekhman, Lior Nezry, Omri Feldman, Natalie Harris, Shashir Reddy, Romina Stella, Ariel Goldstein, Marian Croak, Yossi Matias, Amir Feder
Hugging Face Trending Papers
Sep 17

A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents

The paper examines how large language model (LLM) based graphical user interface (GUI) agents respond to digital nudges. Using a randomized online shopping experiment with 3,600 agents across six frontier models, it finds that agents are vulnerable to both automatic and reflective nudges. The study shows that the agents’ reasoning configuration moderates these effects in opposite directions—reducing susceptibility to automatic nudges while increasing it to reflective social influence nudges—and that this redirection is systematically linked to model scale.