Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjective. We present a resource of 1,482 tweets from politically controversial discourse, annotated by five annotators for the presence of enthymemes and their argument structure, designed to study label variation.
Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are rarely interpretable from argumentative spans alone...
arXiv:2609.10192v1 Announce Type: new
Abstract: Political debates are often analyzed through Argument Mining (AM) to investigate the key arguments that drive them. However, political arguments are ra...
By Lucio La Cava, Stefano Francesco Monea, Sergio Greco
The paper introduces a retrieval‑augmented framework for detecting and classifying fallacies in political debate transcripts. By dynamically retrieving documents guided by argumentative relations of support and attack, the method leverages external knowledge to improve performance. Experiments on the ElecDeb60to20 benchmark show significant gains, raising macro‑F1 to 0.864 for detection and 0.725 for classification compared to non‑retrieval baselines.
By Deborah Dore, Greta Damo, Elena Cabrio, Serena Villata
The paper introduces “Persuasio”, a multi‑agent dialogue platform that uses a formal argumentation theory to adjudicate winners in free‑text debates. Using this system, the authors generated 192 debates on a UK political topic involving humans and large language models (LLMs), and evaluated 22 interlocutors through automated adjudication and 9,702 crowdsourced pairwise judgments across 1,386 annotation instances. The results show a consistent decoupling between subjective persuasiveness—where LLMs dominate—and formal argumentative strength—where humans remain competitive, with multi‑agent and retrieval‑augmented variants widening this gap.
By Jordan Robinson, Angus R. Williams, Katie Atkinson, Anthony G. Cohn
arXiv:2508. 03250v4 Announce Type: replace-cross Abstract: The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens.
By Deborah Dore, Elena Cabrio, Serena Villata
arXiv:2406. 14657v4 Announce Type: replace-cross Abstract: We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
By Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Shwartz-Ziv
MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.
By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv:2603. 23841v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact their objectivity.
By Rohan Khetan, Ashna Khetan
The paper introduces WSF-ARG+, a new dataset that pairs hate speech with check‑worthiness annotations, and presents an LLM‑in‑the‑loop framework to streamline the annotation process. Experiments with 12 open‑weight large language models demonstrate that the framework cuts human effort while maintaining annotation quality. The study also shows that incorporating check‑worthiness labels improves hate‑speech detection performance, boosting macro‑F1 scores for large models by up to 0.213 and averaging 0.154 across models.
By Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin
arXiv:2606. 01736v1 Announce Type: cross Abstract: As LLMs are increasingly used to draft public-facing arguments, they may flatten public debate by repeatedly introducing the same polished, plausible arguments.
By Yekyung Kim, Yapei Chang, Chau Minh Pham, Mohit Iyyer
The paper introduces the Persuasion Index (PI), a taxonomy of 15 persuasion dimensions grounded in psychological and communication theories, implemented with 55 lexicon- and rule-based sub-features. PI is modular, allowing individual detectors to be swapped while preserving its theoretical framework. Evaluations on four English argumentative datasets show that PI provides a shared, lightweight feature space that captures meaningful predictive signals and reveals consistent dimension-level associations with persuasion outcomes, with variations across topics and stances.
By Liancheng Gong, Zhiyang Wang, Yiwei Xu, Julia Mendelsohn