arXiv Computation and Language

How To Do Things With Prompts

The paper examines how users’ prompts to large language models evolve over time, applying speech act and politeness theory to a corpus of 2,000 English prompts from 2023 and 2025. It finds a shift toward more indirect, implicit, and fragmentary directive speech acts, with a notable 14.9‑percentage‑point drop in explicit propositional content and a decline in politeness markers. This suggests users increasingly rely on the model’s inferential abilities, treating it as a competent implicature resolver.

arXiv Computation and Language
Sep 25

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors compare seven LLMs to human annotations. They find that models agree more with each other than with humans, show systematic neutral bias in categorical predictions, and that alignment varies with explicit linguistic cues and rapport‑building strategies.

By Rong Wang, Kun Sun, Yadong Guo
Hugging Face Trending Papers
Sep 24

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors find that LLMs agree more with each other than with humans. They observe that model–human alignment depends on explicit linguistic cues, while misaligned cases often involve rapport‑building strategies. Additionally, models tend to overproduce Neutral labels and underpredict Impolite labels, a pattern that persists even when expert consensus is used as a reference.

arXiv AI
Aug 20

The Deontic Gap: Large Language Models and the Modal Language of Obligation

The study investigates how large language models (LLMs) handle deontic modal verbs such as must, should, and have to, comparing AI-generated text to contemporary human usage across multiple corpora. Results show that LLMs consistently underuse positive deontic modals relative to modern informal digital contexts, though their modal frequencies align with formal published English from the 20th century. The underuse is most pronounced in constructions tied to interpersonal stance, while LLMs match or exceed human usage in instructional or question‑answering contexts, suggesting genre‑dependent modal profiles.

By Daniel Hart, Sarah Allred, Joseph Abbas, Morenike Alugo
arXiv Computation and Language
Sep 1

How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations

The paper investigates how the way users phrase advice‑seeking requests—termed articulation—creates stable, measurable patterns distinct from the topics of the requests. By analyzing 16,447 prompts from public chat corpora, the authors identify a small set of latent articulation factors that consistently appear across datasets and splits. One key finding is a long‑form, information‑poor style that leads language models to give shorter, vaguer answers without seeking clarification, a pattern that persists across topics and prompt lengths.

By Juneha Baek, Suhyeon Lee, Donghyuk Shin
arXiv Computation and Language
Sep 18

Full-Duplex Speech Models Take the Floor When Asked, Not When Needed

Full‑duplex speech models can listen and speak simultaneously, but they struggle to decide when to speak. Experiments with five model families show that being addressed or encountering silence are reliable triggers, whereas cues like false facts or hazards are not. Even when models answer questions, they rarely challenge false claims or warn about danger, revealing a gap in content understanding and intervention decisions.

By Linkai Peng, Baorian Nuchged, Kaiqi Fu, Yuyang Yao