You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
Read the original on arXiv AI →The paper introduces a benchmark for testing large language models (LLMs) on their ability to infer social pragmatic meanings in indirect and playful Chinese online comments. Using over 200,000 public social media interactions, the authors created 4,735 human-validated diagnostic items that pair a target comment with its preceding context and plausible misreadings. Eight LLMs were evaluated in a cross-writer setting, with the best model achieving 81.42% leave-writer-out accuracy, while human accuracy reached 90.8%. The study finds that models can detect broad irony or playfulness but often misidentify the specific mechanism or interactional move.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.