arXiv AI

LLM-based Models for Detecting Emerging Topics in Service Feedback

arXiv:2606. 26595v1 Announce Type: new Abstract: Enhancing the analysis of service feedback is essential for public sector organizations, particularly tax administrations, where trust and compliance depend on fair and effective service delivery.

arXiv Computation and Language
Sep 22

LLJ Cards: Best practices for the Use of LLMs as Judges

arXiv:2609.24516v1 Announce Type: new Abstract: In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these s...

By Khaoula Chehbouni, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, Jackie Chi Kit Cheung
arXiv AI
6d ago

A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models

The article surveys fake review detection research, focusing on how pre‑trained language models (PLMs) and large language models (LLMs) influence both the generation of deceptive reviews and their detection. It reviews 211 studies from 2018 to early 2026, categorizing methods by evidence source—such as review text, sentiment, rating behavior, temporal metadata, user‑product graphs, multimodal content, external knowledge, and LLM‑generated signals—and by fusion level. The survey traces the evolution from traditional machine learning to PLM‑based and LLM‑based approaches, evaluates performance on Amazon, Yelp, and OpSpam benchmarks, and highlights open challenges including adversarial generation, cross‑domain transfer, uncertainty‑aware fusion, robustness to missing sources, interpretability, and trustworthy evaluation of AI‑generated deceptive content.

By Fanji Yang (Guizhou University of Finance and Economics), Huiyao Chen (Harbin Institute of Technology), Xi Yu (Guizhou University of Finance and Economics), Meishan Zhang (Harbin Institute of Technology), Xiaohong Xiao (Guizhou University of Commerce), Mingsen Deng (Guizhou University of Finance and Economics)
arXiv AI
Sep 25

From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring

The paper explores using large language models (LLMs) as AI respondents to convert policy documents into structured survey responses. It introduces a long-context in‑context learning pipeline that maps policy text to predefined survey categories such as policy instruments, target groups, and thematic areas, and includes a secondary LLM validation step. Evaluation on a multi‑country dataset shows high agreement (84‑95%) with human responses for structured indicators, though free‑text fields differ, indicating that hybrid human‑AI workflows can enhance policy monitoring efficiency while still requiring human oversight.

By Carolyn Cole, Matthias Deschryvere, Toqeer Ehsan, Arash Hajikhani
Hugging Face Trending Papers
Aug 6

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency.

arXiv AI
Sep 11

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

XAI-Arena proposes using large language models (LLMs) as judges to evaluate the quality of explainable AI (XAI) explanations, aiming for reproducibility, scalability, and multidimensional assessment. The framework assesses dimensions such as simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability across different datasets, models, and stakeholder personas. Human validation shows a strong positive correlation between LLM-generated and human ratings (Spearman's rho = .693, p < .001), supporting the viability of LLM-based evaluations.

By Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel
arXiv AI
Jun 6

SAGE: Scalable AI Governance & Evaluation

arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.

By Benjamin Le, Xueying Lu, Nick Stern, Wenqiong Liu, Igor Lapchuk, Xiang Li, Baofen Zheng, Kevin Rosenberg, Jiewen Huang, Zhe Zhang, Abraham Cabangbang, Satej Milind Wagle, Jianqiang Shen, Raghavan Muthuregunathan, Abhinav Gupta, Mathew Teoh, Andrew Kirk, Thomas Kwan, Jingwei Wu, Wenjing Zhang
arXiv Computation and Language
Aug 27

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

IDEAlign introduces a new protocol for evaluating the similarity of large language model (LLM) annotations to expert judgments. It uses pick‑the‑odd‑one‑out tasks to capture expert similarity and benchmarks various similarity methods—including text embeddings, topic models, and LLM-as-a-judge—against these human ratings. Applied to educational datasets, the study finds that most metrics miss nuanced expert dimensions, with LLM-as-a-judge performing best yet still insufficient for full expert alignment.

By Hyunji Nam, Lucia Langlois, James Malamut, Mei Tan, Dorottya Demszky
arXiv Computation and Language
Sep 18

What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

The study analyzes 17,012 app‑store reviews for six major generative‑AI apps, using BERTopic and RoBERTa to uncover topics and sentiment. Negative sentiment is most common around advertising, authentication, server reliability, and subscription pricing, with significant differences across apps—Claude shows the highest negative sentiment yet a highly enthusiastic user base. The authors also note geopolitical and privacy concerns for DeepSeek and propose a Trust Friction Score to quantify trust and usability barriers.

By Md Jafrin Hossain, Umme Nusrat Jahan, Shouvaggo Sharif Shammo