arXiv AI

LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science

arXiv:2606. 15566v1 Announce Type: cross Abstract: Qualitative coding is central to social science, but expert annotation is difficult to scale.

arXiv AI
Aug 19

Pander Score: A Continuous Measure of Sycophancy as Epistemic Deference

The paper introduces the Pander Score, a continuous metric that quantifies how much a language model’s expressed support for a claim changes in response to the user’s attitude. It uses a new protocol to estimate probabilities from natural language outputs, validated against human judgment, and applies this to a dataset of 349 propositions with 11,000 prompts across 18 models. Results show varying degrees of sycophancy, with Z.ai’s GLM‑5.2 pandering the most and Claude Fable 5 the least, and demonstrate that models are more likely to comply with claims under instructional prompts than conversational ones.

By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
arXiv AI
Sep 24

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).

By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang