arXiv:2604. 12138v2 Announce Type: replace Abstract: This position paper argues that Retrieval-Augmented Generation systems exhibit a systematic factual bias-optimizing for epistemic uncertainty reduction while ignoring the aleatoric uncertainty inherent in opinion-rich content - and that this misalignment demands a paradigm shift in retrieval system design.
By Aditya Agrawal, Alwarappan Nakkiran, Darshan Fofadiya, Alex Karlsson, Harsha Aduri
arXiv:2606. 16845v1 Announce Type: cross Abstract: Large Language Models (LLMs) natively default to literal semantic interpretations, making zero-shot irony detection a persistent challenge.
By Ankit Bhattacharjee, Krityapriya Bhaumik
arXiv:2609.36194v1 Announce Type: new
Abstract: Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducib...
By Muhammad Abdullahi Said, Abass Oguntade, Elisha Komolafe, Babangida Sani, Fatima Muhammad Adam, Muhammad Sammani Sani
arXiv:2608. 15619v1 Announce Type: new Abstract: Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipline.
By Keito Inoshita
The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.
By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.
By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
arXiv:2601. 05232v3 Announce Type: replace-cross Abstract: Most people now get their news from videos on social media, such as YouTube and Facebook, rather than through curated journalism.
By P. Gilda (Columbia University), P. Dungarwal (Columbia University), A. Thongkham (Columbia University), E. T. Ajayi (St John's University), S. Choudhary (Columbia University), T. M. Terol (Columbia University), C. Lam (Columbia University), J. P. Araujo (Columbia University), M. McFadyen-Mungalln (Columbia University), L. S. Liebovitch (Columbia University), P. T. Coleman (Columbia University), H. West (Columbia University), K. Sieck (Toyota Research Institute), S. Carter (Toyota Research Institute)
arXiv:2609.24574v1 Announce Type: new
Abstract: Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the...
By Hazem Ibrahim, Yasir Zaki
arXiv:2508.13533v2 Announce Type: replace
Abstract: Within a model family, a smaller variant is often deployed as a drop-in replacement for a larger one when their performance is similar. However, pe...
By Rohit Raj Rai, Chirag Kothari, Siddhesh Shelke, Yatika Jena, Amit Awekar
arXiv:2605.01017v3 Announce Type: replace
Abstract: We introduce Xiaohongshu Social Comparison Reader Elicitation (XHS-SCoRE), a reader-grounded benchmark for detecting whether text-only Xiaohongshu...
By Hua Zhao, Jiapei Gu, Michelle Mingyue Gu
The paper introduces Label-Confidence-Aware Uncertainty Quantification (LCA-UQ), a method that uses Pointwise Kullback-Leibler divergence to align global entropy from multiple stochastic samples with the local confidence of a candidate answer. By bridging this gap, LCA-UQ improves the reliability and stability of uncertainty assessments in natural language generation. Experiments on popular LLMs and NLP datasets show that label sources significantly influence classification and that LCA-UQ outperforms existing uncertainty estimation approaches.
By Qinhong Lin, Yinglun Feng, Yuhao Zhang, Zhongliang Yang, Linna Zhou
Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extremely sparse outputs.