When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak labels for textual sentiment, yet whether the two actually agree is rarely questioned.
arXiv:2601. 05261v2 Announce Type: replace-cross Abstract: Online consumer reviews are important decision-support resources in e-commerce, yet the increasing volume of reviews often creates information overload and makes it difficult for users to identify content that matches their individual preferences.
By Muhammad Jawad Mufti, Omar Hammad, MD. Mahfuzur Rahman
The study analyzes 17,012 app‑store reviews for six major generative‑AI apps, using BERTopic and RoBERTa to uncover topics and sentiment. Negative sentiment is most common around advertising, authentication, server reliability, and subscription pricing, with significant differences across apps—Claude shows the highest negative sentiment yet a highly enthusiastic user base. The authors also note geopolitical and privacy concerns for DeepSeek and propose a Trust Friction Score to quantify trust and usability barriers.
By Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo
arXiv:2609.40241v1 Announce Type: cross
Abstract: Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation q...
By Hanjia Lyu, Yinglong Xia
The paper introduces Vocalizer, a mobile app that lets users submit spoken online reviews enhanced by a large language model. A longitudinal study shows that users often use the AI to add detail and that interactive AI features boost confidence and willingness to share reviews. The authors also outline the benefits and challenges of embedding AI assistance in review-writing systems.
By Kavindu Perera, D\'aniel Szab\'o, Niels van Berkel, Aku Visuri, Chi-Lan Yang, Koji Yatani, Simo Hosio
arXiv:2608.22859v1 Announce Type: cross
Abstract: RAG systems are increasingly used to summarize what large collections of documents say. A user asks "What do people think about X?" and receives an a...
By Aman Singh Thakur, Aditya Agrawal, Alwarappan Nakkiran, Alex Karlsson
This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.
By Noah Mami\'e, Laurin van den Bergh
The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.
arXiv:2504. 14053v2 Announce Type: replace-cross Abstract: Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the number that is supposed to separate good listings from bad ones barely varies.
By Ali Safari
arXiv:2601. 21817v2 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm.
By Mingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou Zhou
arXiv:2606. 13858v1 Announce Type: cross Abstract: Recommendation systems are essential in modern music streaming platforms due to the vast amount of available content.
By Terence Zeng, Abhishek K. Umrawal
The paper introduces SARA, an industrial framework that scales articulated user rationales (AURs) for recommendation systems. It curates a high‑quality AUR dataset from 240 M users, trains a 7B‑parameter MLLM (SARA‑7B) to generate rationales for millions of authors, and integrates these generated rationales into a production ranking model (SARA‑Ranker). Offline and online experiments demonstrate that the system produces more specific, polarity‑consistent rationales and improves user engagement while reducing negative feedback.
By Haoke Xiao, Yueyang Liu, Yuhui Zhang, Xiang Chen, Yufei Liu, Jia Xu, Yalong Guan, Xiaolan Zhu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zijie Meng, Zejian Zhang, Ruochen Yang, Xiangyu Wu, Tingting Gao, Han Li, Lantao Hu, Cheng Luo, Kun Gai