arXiv Computation and Language By Neha Srikanth, Jordan Boyd-Graber, Rachel Rudinger

DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering

Read the original on arXiv Computation and Language →

DiscoTrace is a method that identifies rhetorical strategies used by answerers to information‑seeking questions by representing answers as sequences of question‑related discourse acts paired with interpretations of the original question, annotated on top of rhetorical structure theory parses. When applied to answers from nine different communities, DiscoTrace reveals that these communities exhibit diverse preferences for answer construction, whereas large language models (LLMs) lack such rhetorical diversity even when prompted to follow specific community guidelines. Additionally, LLMs tend to adopt a breadth‑oriented approach, addressing interpretations of questions that human answerers often ignore, highlighting a systematic difference in how LLMs and humans respond to information needs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
2d ago

Rhetorical Questions in LLM Representations: A Linear Probing Study

The study investigates how large language models encode rhetorical questions by applying linear probes to two social‑media datasets. It finds that rhetorical signals appear early in the model’s representations, are most stable in last‑token embeddings, and can be distinguished from information‑seeking questions with AUROC 0.7–0.8 even across datasets. However, probes trained on different datasets rank target instances differently, revealing that multiple, distinct linear directions capture various rhetorical cues rather than a single shared representation.

By Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
arXiv Computation and Language
Sep 25

Likelihood Ranking doesn't Scale Like Prompting in LLMs

The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.

By Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau
arXiv Computation and Language
Sep 1

Evaluating the Capabilities of LLMs for Persuasive Dialogue

The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.

By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn