arXiv AI

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

The paper introduces a benchmark for testing large language models (LLMs) on their ability to infer social pragmatic meanings in indirect and playful Chinese online comments. Using over 200,000 public social media interactions, the authors created 4,735 human-validated diagnostic items that pair a target comment with its preceding context and plausible misreadings. Eight LLMs were evaluated in a cross-writer setting, with the best model achieving 81.42% leave-writer-out accuracy, while human accuracy reached 90.8%. The study finds that models can detect broad irony or playfulness but often misidentify the specific mechanism or interactional move.

arXiv Computation and Language
Sep 4

Benchmarking Machine Translation on Chinese Social Media Texts

The paper introduces CSM-MTBench, a benchmark for evaluating machine translation on Chinese social media text. It addresses two main challenges: limited parallel data due to slang and stylistic nuances, and inadequate evaluation metrics that miss these informal features. The benchmark includes two expert-curated subsets—Fun Posts and Social Snippets—and proposes specialized evaluation methods for each, revealing significant differences among over 20 MT models in handling semantic and stylistic aspects.

By Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu, Yao Hu, Shaosheng Cao
arXiv Computation and Language
Sep 1

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

The study investigates how multimodal large language models (MLLMs) use prosodic cues in sarcasm detection. By testing Qwen2.5‑Omni and Qwen3‑Omni on Mandarin Chinese and English across five modality conditions, the authors find that adding audio increases false positives without improving true positives. Acoustic error analysis shows that models rely on a stereotypical prosodic pattern—elevated pitch and irregular pausing—that does not align with genuine sarcasm cues, and manipulating these dimensions alone can raise false positive rates up to 60%. The same effect appears in Gemini 3 Flash Preview, indicating the heuristic is not limited to a single architecture.

By Yongjian Chen, Pengfei Wei, Yiqun Sun, Zhu Li, Lawrence B. Hsieh
arXiv Computation and Language
Sep 1

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.

By Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce