arXiv Computation and Language

Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing

The paper presents a statistically rigorous sentiment index for Google Play user reviews, combining normalized star ratings and text-sentiment scores through covariance-aware inverse-variance weighting. It aggregates review-level estimates using bounded helpfulness and recency weights, then applies Gaussian-conjugate shrinkage toward a population mean based on estimated precision. The authors also provide distributional diagnostics for different API sort orders, avoid inappropriate Kolmogorov‑Smirnov tests for discrete data, and use a Kalman filter to smooth temporal trends, all supported by full mathematical proofs.

arXiv Computation and Language
Sep 18

What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

The study analyzes 17,012 app‑store reviews for six major generative‑AI apps, using BERTopic and RoBERTa to uncover topics and sentiment. Negative sentiment is most common around advertising, authentication, server reliability, and subscription pricing, with significant differences across apps—Claude shows the highest negative sentiment yet a highly enthusiastic user base. The authors also note geopolitical and privacy concerns for DeepSeek and propose a Trust Friction Score to quantify trust and usability barriers.

By Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo
arXiv AI
Sep 15

From Voice to Value: Leveraging AI to Enhance Spoken Online Reviews on the Go

The paper introduces Vocalizer, a mobile app that lets users submit spoken online reviews enhanced by a large language model. A longitudinal study shows that users often use the AI to add detail and that interactive AI features boost confidence and willingness to share reviews. The authors also outline the benefits and challenges of embedding AI assistance in review-writing systems.

By Kavindu Perera, D\'aniel Szab\'o, Niels van Berkel, Aku Visuri, Chi-Lan Yang, Koji Yatani, Simo Hosio
arXiv Computation and Language
Aug 25

WARP: Wasserstein-Aligned RAG for Population Opinions

arXiv:2608.22859v1 Announce Type: cross Abstract: RAG systems are increasingly used to summarize what large collections of documents say. A user asks "What do people think about X?" and receives an a...

By Aman Singh Thakur, Aditya Agrawal, Alwarappan Nakkiran, Alex Karlsson
arXiv AI
Sep 18

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.

By Noah Mami\'e, Laurin van den Bergh
Hugging Face Trending Papers
Sep 17

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.

arXiv AI
Sep 17

Scaling Articulated Rationales for MLLM-based Recommendation

The paper introduces SARA, an industrial framework that scales articulated user rationales (AURs) for recommendation systems. It curates a high‑quality AUR dataset from 240 M users, trains a 7B‑parameter MLLM (SARA‑7B) to generate rationales for millions of authors, and integrates these generated rationales into a production ranking model (SARA‑Ranker). Offline and online experiments demonstrate that the system produces more specific, polarity‑consistent rationales and improves user engagement while reducing negative feedback.

By Haoke Xiao, Yueyang Liu, Yuhui Zhang, Xiang Chen, Yufei Liu, Jia Xu, Yalong Guan, Xiaolan Zhu, Xiaoyu Zhang, Shijun Wang, Shuang Yang, Zijie Meng, Zejian Zhang, Ruochen Yang, Xiangyu Wu, Tingting Gao, Han Li, Lantao Hu, Cheng Luo, Kun Gai