arXiv AI

Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses

arXiv:2606. 01845v1 Announce Type: cross Abstract: Although large language models (LLMs) have shown considerable progress in pragmatic language understanding, prior research has focused mainly on their comprehension of verbal behavior.

Hugging Face Trending Papers
Jul 20

Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification.

arXiv Computation and Language
Sep 25

How To Do Things With Prompts

The paper examines how users’ prompts to large language models evolve over time, applying speech act and politeness theory to a corpus of 2,000 English prompts from 2023 and 2025. It finds a shift toward more indirect, implicit, and fragmentary directive speech acts, with a notable 14.9‑percentage‑point drop in explicit propositional content and a decline in politeness markers. This suggests users increasingly rely on the model’s inferential abilities, treating it as a competent implicature resolver.

By Kristina \v{S}ekrst, Virna Karli\'c
arXiv AI
Aug 26

When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

The paper investigates the role of minimal responses—short, empathic utterances—in psychological counseling, noting that such brief replies are common in human dialogues but underrepresented in large language model (LLM) outputs. Using a two‑stage filtering approach and contextual verification with an LLM, the authors systematically analyze minimal responses across multiple counseling datasets. They find that while strong commercial LLMs can produce minimal replies when prompted, they often fail to judge when these replies are appropriate, and counseling‑specific models trained on synthetic data tend to generate longer, content‑rich responses instead.

By Zhiyang Qi
arXiv AI
Sep 7

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

The paper introduces a benchmark for testing large language models (LLMs) on their ability to infer social pragmatic meanings in indirect and playful Chinese online comments. Using over 200,000 public social media interactions, the authors created 4,735 human-validated diagnostic items that pair a target comment with its preceding context and plausible misreadings. Eight LLMs were evaluated in a cross-writer setting, with the best model achieving 81.42% leave-writer-out accuracy, while human accuracy reached 90.8%. The study finds that models can detect broad irony or playfulness but often misidentify the specific mechanism or interactional move.

By Shiwei Hong, Junjie Ma, Emma Jiren Wang, Ethan Z. Rong, Siying Hu, Haichang Li, Ziying Wang, Zhicong Lu
arXiv AI
Jul 1

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

arXiv:2604. 22027v2 Announce Type: replace-cross Abstract: One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed.

By Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez, Eric Todd, David Bau, Michael L. Littman, Stephen H. Bach, Ellie Pavlick