arXiv:2606. 00084v1 Announce Type: cross Abstract: Online travel platforms generate vast volumes of user-generated hotel reviews, offering rich opportunities to understand traveler experiences at scale.
By Dineth Jayakody, Pasindu Thenahandi, Sampath Jayarathna
arXiv:2609.36194v1 Announce Type: new
Abstract: Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducib...
By Muhammad Abdullahi Said, Abass Oguntade, Elisha Komolafe, Babangida Sani, Fatima Muhammad Adam, Muhammad Sammani Sani
arXiv:2607. 10825v1 Announce Type: cross Abstract: Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences, and concerns.
By Fabrizio Marozzo, Stefano Iannicelli
The paper presents a statistically rigorous sentiment index for Google Play user reviews, combining normalized star ratings and text-sentiment scores through covariance-aware inverse-variance weighting. It aggregates review-level estimates using bounded helpfulness and recency weights, then applies Gaussian-conjugate shrinkage toward a population mean based on estimated precision. The authors also provide distributional diagnostics for different API sort orders, avoid inappropriate Kolmogorov‑Smirnov tests for discrete data, and use a Kalman filter to smooth temporal trends, all supported by full mathematical proofs.
By Marco Mandap
arXiv:2606. 29614v1 Announce Type: cross Abstract: This study examines whether supervised fine-tuning remains necessary for Turkish sentiment analysis in the era of large language models.
By Sercan Karaka\c{s}, Yusuf \c{S}im\c{s}ek
arXiv:2504. 14053v2 Announce Type: replace-cross Abstract: Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the number that is supposed to separate good listings from bad ones barely varies.
By Ali Safari
The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.
By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
The study analyzes 17,012 app‑store reviews for six major generative‑AI apps, using BERTopic and RoBERTa to uncover topics and sentiment. Negative sentiment is most common around advertising, authentication, server reliability, and subscription pricing, with significant differences across apps—Claude shows the highest negative sentiment yet a highly enthusiastic user base. The authors also note geopolitical and privacy concerns for DeepSeek and propose a Trust Friction Score to quantify trust and usability barriers.
By Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo
arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.
By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv:2607. 05761v1 Announce Type: new Abstract: Modern data-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time-consuming, and difficult to scale.
By Stephen L. France, Pia. A. Albinsson
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.
By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani