arXiv AI

Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

arXiv:2606. 23701v1 Announce Type: cross Abstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure.

arXiv Computation and Language
Sep 18

What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

The study analyzes 17,012 app‑store reviews for six major generative‑AI apps, using BERTopic and RoBERTa to uncover topics and sentiment. Negative sentiment is most common around advertising, authentication, server reliability, and subscription pricing, with significant differences across apps—Claude shows the highest negative sentiment yet a highly enthusiastic user base. The authors also note geopolitical and privacy concerns for DeepSeek and propose a Trust Friction Score to quantify trust and usability barriers.

By Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo
arXiv Machine Learning
Sep 24

From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026

This scoping review examines 421 studies (2015‑2026) on natural language processing applied to student evaluation of teaching comments. It maps the technical evolution from lexicons and classifiers to transformers and large language models, and evaluates four value dimensions. The review identifies a significant gap between actionable outputs (61.3%) and intended‑user evaluation (11.6%), highlighting limited progress in educational value and robustness.

By Jeff Eicher, Rafael da Silva
arXiv AI
Sep 25

An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection

The paper presents an explainable hate‑speech detection framework that combines DistilBERT embeddings, a Bi‑LSTM network, and an attention mechanism to capture contextual and sequential information. It uses LIME to highlight influential text features, providing transparency in predictions. Evaluated on two benchmark datasets for both binary and multi‑class tasks, the model achieves F1‑scores of 96.78%–99.53% for binary classification and 94.99%–97.00% for multi‑class classification, outperforming existing baselines.

By Rameesha Zia, Muhammad Shahid Iqbal Malik