arXiv:2606.12186v2 Announce Type: replace
Abstract: Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjec...
By Martial Pastor, Nelleke Oostdijk
Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are rarely interpretable from argumentative spans alone...
The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.
By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
The paper introduces DNE‑ElecDeb, an enriched version of the USElecDeb dataset that annotates Debate Named Entities (DNEs) in both argumentative and non‑argumentative spans, and defines Debate Named Entity Recognition (DNER) as a new task. It proposes Joint Argument and Entity Tagging (JAET), a generative framework that fine‑tunes decoder‑only LLMs to insert inline argument and entity tags into debate turns while preserving the original transcript. JAET achieves significant improvements in joint AM+DNER performance (+27.3% relative F1 in the untyped setting and +41.9% in the typed setting) over sequential pipelines, and these gains generalize to Persuasive Essays (+26.6% and +52.7%).
By Lucio La Cava, Stefano Francesco Monea, Sergio Greco
The paper introduces a retrieval‑augmented framework for detecting and classifying fallacies in political debate transcripts. By dynamically retrieving documents guided by argumentative relations of support and attack, the method leverages external knowledge to improve performance. Experiments on the ElecDeb60to20 benchmark show significant gains, raising macro‑F1 to 0.864 for detection and 0.725 for classification compared to non‑retrieval baselines.
By Deborah Dore, Greta Damo, Elena Cabrio, Serena Villata
arXiv:2508. 03250v4 Announce Type: replace-cross Abstract: The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens.
By Deborah Dore, Elena Cabrio, Serena Villata
arXiv:2609.27811v1 Announce Type: cross
Abstract: Online platforms have become arenas for the public contestation of climate change, shaping how scientific knowledge, denial, and uncertainty are expr...
By Daniel Morais, Diego H. M. Magalhaes, Gabriel H. Silva, Andrea Failla, Valeria de C. Santos, Helen C. S. C. Lima, Carlos H. G. Ferreira
arXiv:2609.38256v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to answer questions about politically contentious issues, yet evaluations typically treat a model's...
By Olivia Macmillan-Scott, Michael Jacobs, Nils Metternich, Mirco Musolesi
arXiv:2406. 14657v4 Announce Type: replace-cross Abstract: We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
By Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Shwartz-Ziv
The paper investigates how well large language models (LLMs) can handle character attacks—ad hominem arguments—in political debates. By analyzing natural political dialogues and comparing LLM-generated responses to a corpus of U.S. presidential debates, the study finds that most LLMs favor logical defenses and rarely use ethos-based counterattacks. The authors suggest that safety fine‑tuning limits LLMs’ strategic options, preventing them from fully engaging in realistic political discourse.
By Ewelina Gajewska, Katarzyna Budzynska, Jaroslaw Chudziak
MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.
By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv:2609.22133v1 Announce Type: new
Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at...
By Kentaro Nakamura, Jing Ling Tan, George Yean