arXiv Computation and Language By Kristina \v{S}ekrst, Virna Karli\'c

How To Do Things With Prompts

Read the original on arXiv Computation and Language →

The paper examines how users’ prompts to large language models evolve over time, applying speech act and politeness theory to a corpus of 2,000 English prompts from 2023 and 2025. It finds a shift toward more indirect, implicit, and fragmentary directive speech acts, with a notable 14.9‑percentage‑point drop in explicit propositional content and a decline in politeness markers. This suggests users increasingly rely on the model’s inferential abilities, treating it as a competent implicature resolver.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 25

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors compare seven LLMs to human annotations. They find that models agree more with each other than with humans, show systematic neutral bias in categorical predictions, and that alignment varies with explicit linguistic cues and rapport‑building strategies.

By Rong Wang, Kun Sun, Yadong Guo
Hugging Face Trending Papers
Sep 24

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors find that LLMs agree more with each other than with humans. They observe that model–human alignment depends on explicit linguistic cues, while misaligned cases often involve rapport‑building strategies. Additionally, models tend to overproduce Neutral labels and underpredict Impolite labels, a pattern that persists even when expert consensus is used as a reference.