arXiv Computation and Language

DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering

DiscoTrace is a method that identifies rhetorical strategies used by answerers to information‑seeking questions by representing answers as sequences of question‑related discourse acts paired with interpretations of the original question, annotated on top of rhetorical structure theory parses. When applied to answers from nine different communities, DiscoTrace reveals that these communities exhibit diverse preferences for answer construction, whereas large language models (LLMs) lack such rhetorical diversity even when prompted to follow specific community guidelines. Additionally, LLMs tend to adopt a breadth‑oriented approach, addressing interpretations of questions that human answerers often ignore, highlighting a systematic difference in how LLMs and humans respond to information needs.

arXiv AI
2d ago

Rhetorical Questions in LLM Representations: A Linear Probing Study

The study investigates how large language models encode rhetorical questions by applying linear probes to two social‑media datasets. It finds that rhetorical signals appear early in the model’s representations, are most stable in last‑token embeddings, and can be distinguished from information‑seeking questions with AUROC 0.7–0.8 even across datasets. However, probes trained on different datasets rank target instances differently, revealing that multiple, distinct linear directions capture various rhetorical cues rather than a single shared representation.

By Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
arXiv Computation and Language
Sep 25

Likelihood Ranking doesn't Scale Like Prompting in LLMs

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv Computation and Language
Sep 1

Evaluating the Capabilities of LLMs for Persuasive Dialogue

The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.

By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
arXiv Computation and Language
Sep 7

On Epistemic Diversity in Large Language Models

The paper introduces the concept of epistemic diversity for large language models (LLMs), defining it as the range of valid answers, explanations, and reasoning routes that an LLM presents to users. It argues that evaluating LLMs solely on accuracy or alignment is insufficient, especially when LLMs are used for knowledge-intensive tasks. The authors propose a preliminary framework for measuring epistemic diversity and demonstrate that leading LLMs often collapse large valid answer spaces into small canonical subsets, indicating epistemic narrowness.

By Elisabeth Kirsten, Nicole Kr\"amer, Muhammad Bilal Zafar
arXiv AI
Jun 12

Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities

arXiv:2606. 13397v1 Announce Type: cross Abstract: Language operates as a mechanism of both marginalization and resistance, especially for minority communities navigating insensitive and harmful speech online.

By Dipto Das, Achhiya Sultana, Ankit Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed