arXiv AI

Reducing belief in conspiracy theories as they unfold using large language models

arXiv:2608. 06151v1 Announce Type: cross Abstract: The emergence of conspiracy theories in the wake of major events is a significant societal challenge.

arXiv Computation and Language
Sep 7

Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Tendencies in Large Language Models

The paper examines whether large language models (LLMs) exhibit conspiratorial tendencies, socio-demographic biases in this domain, and how easily they can be conditioned to adopt conspiratorial viewpoints. Using validated psychometric surveys, the authors find that LLMs partially align with conspiracy beliefs, that conditioning with demographic attributes yields uneven effects revealing latent biases, and that targeted prompts can readily shift responses toward conspiratorial stances. These findings underscore the vulnerability of LLMs to manipulation and the potential risks of deploying them in sensitive contexts.

By Francesco Corso, Francesco Pierri, Gianmarco De Francisci Morales
arXiv Machine Learning
Sep 25

Agentic Detection of Online Conspiracies

The paper presents an agentic framework for detecting conspiratorial content in social media by inferring the speaker’s intent rather than merely identifying explicit claims. It leverages social context and adaptive tool use, demonstrating superior performance over text-only and non-agentic models on a large Hebrew tweet dataset spanning election cycles and the COVID pandemic. The study highlights the importance of context-aware, reasoning-driven approaches for accurate conspiracy detection.

By Lior Biton, Oren Tsur
arXiv Computation and Language
Aug 31

Beyond the Rabbit Hole: Mapping the Relational Harms of QAnon Radicalization

The paper examines the relational harms of QAnon radicalization by analyzing 12,747 stories from the r/QAnonCasualties support group. Using a computational pipeline, the authors extract thematic traits, cluster them into six radicalization personas, and link these personas to specific emotional harms through LLM-assisted emotion detection and regression modeling. The study finds that certain personas predict distinct emotional outcomes, such as anger and disgust for ideologically driven radicalization, and fear and sadness for personal and cognitive collapse.

By Bich Ngoc Doan, Gianmarco De Francisci Morales, Giuseppe Russo
arXiv AI
3d ago

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.

By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger
arXiv Computation and Language
Sep 25

Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks

The paper investigates how well large language models (LLMs) can handle character attacks—ad hominem arguments—in political debates. By analyzing natural political dialogues and comparing LLM-generated responses to a corpus of U.S. presidential debates, the study finds that most LLMs favor logical defenses and rarely use ethos-based counterattacks. The authors suggest that safety fine‑tuning limits LLMs’ strategic options, preventing them from fully engaging in realistic political discourse.

By Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak