arXiv AI
Sep 7

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

The paper introduces a benchmark for testing large language models (LLMs) on their ability to infer social pragmatic meanings in indirect and playful Chinese online comments. Using over 200,000 public social media interactions, the authors created 4,735 human-validated diagnostic items that pair a target comment with its preceding context and plausible misreadings. Eight LLMs were evaluated in a cross-writer setting, with the best model achieving 81.42% leave-writer-out accuracy, while human accuracy reached 90.8%. The study finds that models can detect broad irony or playfulness but often misidentify the specific mechanism or interactional move.

By Shiwei Hong, Junjie Ma, Emma Jiren Wang, Ethan Z. Rong, Siying Hu, Haichang Li, Ziying Wang, Zhicong Lu
arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
arXiv AI
Sep 3

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses by leveraging weak supervision from emoji distributions. It builds the EmojiDialogue dataset, extending EmpatheticDialogues with emoji votes and confidence scores, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings to generate responses that align with the listener’s stance. In blind pairwise evaluations, EmoStance achieves a 62.2% decisive win rate, notably improving contextual specificity and perceived responsiveness compared to baselines.

By Ziyuan Jin, Yuxuan Ge, Zheng Tian
arXiv Computation and Language
2d ago

Zero-shot narrative detection in social messaging

The paper explores how large language models can detect hidden narratives in social messages without training data. By feeding the models human-written narrative descriptions, performance improves markedly, while automatically generated descriptions or few-shot examples can hurt accuracy. Ensemble techniques, especially majority voting, further boost robustness, and larger models show the best results with less sensitivity to prompts.

By Jes\'us M. Fraile-Hern\'andez, Anselmo Pe\~nas, Patrick Giedemann