arXiv AI

The Deontic Gap: Large Language Models and the Modal Language of Obligation

The study investigates how large language models (LLMs) handle deontic modal verbs such as must, should, and have to, comparing AI-generated text to contemporary human usage across multiple corpora. Results show that LLMs consistently underuse positive deontic modals relative to modern informal digital contexts, though their modal frequencies align with formal published English from the 20th century. The underuse is most pronounced in constructions tied to interpersonal stance, while LLMs match or exceed human usage in instructional or question‑answering contexts, suggesting genre‑dependent modal profiles.

arXiv Computation and Language
Sep 25

How To Do Things With Prompts

The paper examines how users’ prompts to large language models evolve over time, applying speech act and politeness theory to a corpus of 2,000 English prompts from 2023 and 2025. It finds a shift toward more indirect, implicit, and fragmentary directive speech acts, with a notable 14.9‑percentage‑point drop in explicit propositional content and a decline in politeness markers. This suggests users increasingly rely on the model’s inferential abilities, treating it as a competent implicature resolver.

By Kristina \v{S}ekrst, Virna Karli\'c
arXiv AI
Aug 28

How LLMs Distort Our Written Language

Large language models (LLMs) are widely used to assist writing, but this study shows they alter both tone and meaning of human text. A user study found that heavy LLM use increased neutral essays by nearly 70% and made writers feel less creative and less in their voice. Even when prompted to make only grammar edits, LLMs changed the semantic content of essays and produced AI-generated scientific reviews that were less focused on clarity and significance and scored higher on average.

By Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
arXiv AI
Sep 17

MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

MIRAGE is a controlled study that examines how multimodal personal agents use historical evidence when conversation state changes. The study keeps evidence, questions, and scoring constant while varying only the conversation state, then checks if agents can determine answerability, recover the correct source, and answer from it. Results across seven multimodal backbones show distinct failure regimes before and after compaction, heavy reliance on context continuity by open-weight models, and mixed effects of retrieval pressure on source attribution.

By Yu Liu, Wenxiao Zhang, Cheng Hu, Cong Cao, Fangfang Yuan, Xinyu Wang, Jin B. Hong, Yanbing Liu
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv Computation and Language
Sep 25

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors compare seven LLMs to human annotations. They find that models agree more with each other than with humans, show systematic neutral bias in categorical predictions, and that alignment varies with explicit linguistic cues and rapport‑building strategies.

By Rong Wang, Kun Sun, Yadong Guo