The paper investigates how alignment training, specifically reinforcement learning from human feedback (RLHF), affects the internal partisan structure of a large language model. Using a mechanistic case study on Llama 3.1 8B, the authors find that alignment training does not erase the model’s partisan geometry but compresses its variance, producing consistently balanced, non‑partisan outputs. Sparse autoencoder analysis and feature‑level steering experiments reveal that policy‑encoding features become inactive in the aligned model, indicating a causal disconnect rather than structural removal of partisan knowledge.
By Wendy K. Tam
arXiv:2603. 23841v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact their objectivity.
By Rohan Khetan, Ashna Khetan
The paper examines how Preference Inference (PI) models used in large-scale participatory democracy platforms can alter the perceived consensus and minority support by predicting missing votes. It introduces a collective‑centric evaluation framework that assesses whether inferred votes maintain key properties of the overall preference landscape, rather than focusing solely on individual prediction accuracy. Using the largest multilingual dataset to date—four consultations with over 90,000 participants, 1 million votes, and 22 languages—the study finds that models with similar predictive accuracy can differ markedly in how well they preserve the collective structure, underscoring that accuracy alone is insufficient for evaluating PI in democratic contexts.
By Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner, Nazanin Shafiabadi, Laur\`ene Cave, David Mas, Jean-Philippe Cointet, Benjamin Piwowarski, Fran\c{c}ois Yvon
The paper introduces DNE‑ElecDeb, an enriched version of the USElecDeb dataset that annotates Debate Named Entities (DNEs) in both argumentative and non‑argumentative spans, and defines Debate Named Entity Recognition (DNER) as a new task. It proposes Joint Argument and Entity Tagging (JAET), a generative framework that fine‑tunes decoder‑only LLMs to insert inline argument and entity tags into debate turns while preserving the original transcript. JAET achieves significant improvements in joint AM+DNER performance (+27.3% relative F1 in the untyped setting and +41.9% in the typed setting) over sequential pipelines, and these gains generalize to Persuasive Essays (+26.6% and +52.7%).
By Lucio La Cava, Stefano Francesco Monea, Sergio Greco
Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are rarely interpretable from argumentative spans alone...
MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.
By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv:2608. 06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks.
By Debodeep Banerjee, Amitangshu Dasgupta
arXiv:2609.15207v1 Announce Type: new
Abstract: Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before ele...
By Bastiaan Bruinsma, Annika Fred\'en, Paul R\"ottger, Moa Johansson, Asad Sayeed
arXiv:2609.23039v1 Announce Type: new
Abstract: LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their...
By Joan C. Timoneda
The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.
By Ionel Eduard Stan, Paolo Napoletano
arXiv:2601. 19082v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users.
By Trung-Kiet Huynh, Dao-Sy Duy-Minh, Thanh-Bang Cao, Phong-Hao Le, Hong-Dan Nguyen, Phu-Quy Nguyen-Lam, Minh-Luan Nguyen-Vo, Hong-Phat Pham, Phu-Hoa Pham, Thien-Kim Than, Chi-Nguyen Tran, Huy Tran, Gia-Thoai Tran-Le, Alessio Buscemi, Le Hong Trang, The Anh Han
arXiv:2608. 10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems.
By Maurice Flechtner