Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,601 stories · RSS feed

arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv Computation and Language
Sep 2

DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering

DiscoTrace is a method that identifies rhetorical strategies used by answerers to information‑seeking questions by representing answers as sequences of question‑related discourse acts paired with interpretations of the original question, annotated on top of rhetorical structure theory parses. When applied to answers from nine different communities, DiscoTrace reveals that these communities exhibit diverse preferences for answer construction, whereas large language models (LLMs) lack such rhetorical diversity even when prompted to follow specific community guidelines. Additionally, LLMs tend to adopt a breadth‑oriented approach, addressing interpretations of questions that human answerers often ignore, highlighting a systematic difference in how LLMs and humans respond to information needs.

By Neha Srikanth, Jordan Boyd-Graber, Rachel Rudinger
arXiv AI
Sep 2

Visual Attention Faithfulness in Vision-Language Models is Heterogeneous

The study investigates whether attention weights in Vision‑Language Models (VLMs) accurately reflect model reasoning for visual inputs. Using causal perturbation analysis, it identifies three distinct processing modes—Faithful‑Sufficient, Faithful‑Distributed, and Non‑Focal—indicating heterogeneous visual attention faithfulness. The research also shows that human‑annotated ground‑truth regions align with model attention in only about 60% of cases, highlighting a systematic divergence between model visual reliance and human intuition across VQA, document, and chart tasks.

By Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo
arXiv AI
Sep 2

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

The paper introduces WHALE, a method that alternates between updating a language model’s weights and searching for a better harness (the code that manages context and control flow). By iteratively fine‑tuning the model under the current harness and then optimizing the harness under the updated model, WHALE improves performance across search QA, math reasoning, and chess puzzles, outperforming weight‑only, harness‑only, and Fast‑Slow Training by 4.15–24.38 percentage points in mean@8 accuracy. The approach uses either fixed phase lengths or an adaptive patience rule to decide when to switch phases, and the authors provide code on GitHub.

By Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee