The paper introduces a benchmark for testing large language models (LLMs) on their ability to infer social pragmatic meanings in indirect and playful Chinese online comments. Using over 200,000 public social media interactions, the authors created 4,735 human-validated diagnostic items that pair a target comment with its preceding context and plausible misreadings. Eight LLMs were evaluated in a cross-writer setting, with the best model achieving 81.42% leave-writer-out accuracy, while human accuracy reached 90.8%. The study finds that models can detect broad irony or playfulness but often misidentify the specific mechanism or interactional move.
By Shiwei Hong, Junjie Ma, Emma Jiren Wang, Ethan Z. Rong, Siying Hu, Haichang Li, Ziying Wang, Zhicong Lu
The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.
By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses by leveraging weak supervision from emoji distributions. It builds the EmojiDialogue dataset, extending EmpatheticDialogues with emoji votes and confidence scores, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings to generate responses that align with the listener’s stance. In blind pairwise evaluations, EmoStance achieves a 62.2% decisive win rate, notably improving contextual specificity and perceived responsiveness compared to baselines.
By Ziyuan Jin, Yuxuan Ge, Zheng Tian
arXiv:2608.29803v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
By Lin Chen, Yitong Chen, Yong Li
The paper explores how large language models can detect hidden narratives in social messages without training data. By feeding the models human-written narrative descriptions, performance improves markedly, while automatically generated descriptions or few-shot examples can hurt accuracy. Ensemble techniques, especially majority voting, further boost robustness, and larger models show the best results with less sensitivity to prompts.
By Jes\'us M. Fraile-Hern\'andez, Anselmo Pe\~nas, Patrick Giedemann
arXiv:2607. 14957v1 Announce Type: new Abstract: Online firestorms are rapid collective escalations of highly negative user-generated content and may cause substantial reputational and economic damage.
By Besim Shala, Peter Mandl, Andreas Humpe, Martin H\"ausl