arXiv AI

Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates

arXiv:2606. 03029v1 Announce Type: cross Abstract: A core goal of computational social science is to discover interpretable differences in how language varies across outcomes of interest, such as political affiliation or instructional quality.

arXiv Computation and Language
Sep 1

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

The paper examines how large language models (LLMs) respond to different demographic cues—such as names—when users seek advice, focusing on race and gender in a U.S. context. It finds that using different cues for the same group leads to only partially overlapping changes in model responses, producing inconsistent conclusions about personalization and unstable bias metrics. The authors argue that LLMs react to linguistic signals tied to cues rather than to stable demographic categories, and they call for evaluations that use multiple cues and consider underlying mechanisms.

By Manuel Tonneau, Neil K. R. Sehgal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera, Ana Mar\'ia Mu\~noz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, Valentin Hofmann
arXiv Computation and Language
Sep 1

Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

The paper proposes a method to steer large language models (LLMs) to generate explanations tailored to specific target groups. It first identifies group-specific attributes related to explanatory style and knowledge, then uses activation engineering to compute steering vectors that are added to the LLM’s internal activations during inference. Experiments show that this attribute-based steering improves specificity and factuality of explanations compared to prompting and existing steering baselines, and a human study confirms better tailoring to target groups.

By Leandra Fichtel, Janek Prange, Henning Wachsmuth
arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin
arXiv AI
Sep 10

Mapping the Emerging Social Science of Large Language Models

The paper maps the nascent social‑science literature on large language models (LLMs) by analysing 198 curated papers and 47,719 field‑scale papers. It identifies three main domains—LLM as Social Minds, LLM Societies, and LLM‑Human Interactions—each containing 13 subcategories such as reasoning, bias, collective intelligence, and trust. The taxonomy is validated through clustering stability, author classification agreement, and topic mapping, revealing differing prominence across conference and journal venues.

By Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao