arXiv Computation and Language

Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models

The paper examines whether large language models (LLMs) exhibit conspiratorial tendencies, socio-demographic biases in this domain, and how easily they can be conditioned to adopt conspiratorial viewpoints. Using validated psychometric surveys, the authors find that LLMs partially align with conspiracy beliefs, that conditioning with demographic attributes yields uneven effects revealing latent biases, and that targeted prompts can readily shift responses toward conspiratorial stances. These findings underscore the vulnerability of LLMs to manipulation and the potential risks of deploying them in sensitive contexts.

arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv Computation and Language
Sep 17

Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

The paper examines how large language models (LLMs) alter the expression of Dark Triad traits—Machiavellianism, narcissism, and psychopathy—when prompted to fake good or fake bad. Across seven state‑of‑the‑art models and two real‑world contexts (employment selection and forensic evaluation), most models lowered trait scores under fake‑good conditions and raised them under fake‑bad conditions, with varying consistency across traits and models. The study also finds that explicit fake‑bad instructions produce stronger distortions than contextual framing alone, underscoring the influence of motivational and situational context on personality‑related outputs.

By Victoria Popa, Guglielmo Cola, Caterina Senette, Maurizio Tesconi
arXiv Computation and Language
Aug 31

Beyond the Rabbit Hole: Mapping the Relational Harms of QAnon Radicalization

The paper examines the relational harms of QAnon radicalization by analyzing 12,747 stories from the r/QAnonCasualties support group. Using a computational pipeline, the authors extract thematic traits, cluster them into six radicalization personas, and link these personas to specific emotional harms through LLM-assisted emotion detection and regression modeling. The study finds that certain personas predict distinct emotional outcomes, such as anger and disgust for ideologically driven radicalization, and fear and sadness for personal and cognitive collapse.

By Bich Ngoc Doan, Gianmarco De Francisci Morales, Giuseppe Russo
arXiv AI
Sep 10

Beliefs and Behavior in Language Models

arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...

By Alex Smolin, Bryan Wilder
arXiv AI
Jul 24

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

arXiv:2607. 20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed systems.

By Adam Rigby, Raz Saremi, Azadeh Sohrabinejad, Mehdi Rahimi