arXiv AI

Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare

arXiv:2606. 26104v1 Announce Type: cross Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare.

arXiv AI
Sep 24

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

The study examines how different editorial framings in prompts influence large language models’ statistical analysis reports. Using a 4×4 factorial design, researchers found that certain framings—particularly brutally critical prompts on genuine effects and significance-seeking prompts on underpowered nulls—led to factual misrepresentations. Tone shifts were more widespread, with critical framing inducing defensive language across all data patterns, while a confound in the data largely prevented both factual and tonal distortions.

By Paras Balani, Subhrakanta Panda
arXiv Computation and Language
Sep 2

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.

By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv AI
Jun 4

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605. 16301v2 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries.

By Isabella Luong, Joyee Chen, Arturs Kanepajs, Jasmine Brazilek, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu
Hugging Face Trending Papers
Sep 17

When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

The paper investigates whether language‑grounded explanations improve trust in automated sheep pain detection from facial expressions. It finds that attention‑based explanations tied to the Sheep Pain Facial Expression Scale (SPFES) are largely ineffective, prompting the authors to replace the appearance bypass with a concept bottleneck that uses SPFES concept scores supervised by per‑region annotations. This new architecture slightly reduces performance but yields more semantically meaningful concepts, recovering minority pain states and revealing ear‑and‑eye severity orderings without explicit severity supervision, and highlights the pitfalls of pooled concept accuracy under clinical imbalance.

arXiv Computation and Language
Sep 18

Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

The paper introduces the concept of summarization bias in large language models (LLMs), describing a systematic tendency for LLMs to represent narrative meaning as an abstract summary label rather than the reconstructable inferential structure that produces it. It frames this bias within the Bulut Doctrine’s told‑shown axis, arguing that LLMs fail in a specific direction: they default to told‑mode explicitness in generative tasks and reward told‑mode explicitness while under‑detecting shown‑mode suppression in evaluative tasks. The authors outline two regimes of bias, present preliminary evidence, and pre‑register a test protocol to validate or abandon the construct.

By Levent Bulut