arXiv Machine Learning

LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade

The paper evaluates large language models (LLMs) for annotating German parliamentary debates on migration, achieving macro‑F1 scores comparable to human agreement, particularly with GPT‑5 and gpt‑oss‑120B. It combines soft‑label outputs with Design‑based Supervised Learning to mitigate systematic bias and applies the method to a 150‑plus‑year corpus, revealing high solidarity post‑war and a sharp rise in anti‑solidarity since 2015. The study demonstrates that LLMs can enable large‑scale social‑scientific analysis while highlighting the need for rigorous validation and bias correction.

arXiv Machine Learning
Aug 18

Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles

arXiv:2608. 15871v1 Announce Type: cross Abstract: Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour.

By Roman Neruda, Martin Bako\v{s}, Josef \v{S}lerka, V\'it Tu\v{c}ek, Petra Vidnerov\'a, Gabriela Kadlecov\'a
arXiv Computation and Language
Sep 10

Who Argues What? Joint Argument-Entity Detection and Classification in Political Debates

The paper introduces DNE‑ElecDeb, an enriched version of the USElecDeb dataset that annotates Debate Named Entities (DNEs) in both argumentative and non‑argumentative spans, and defines Debate Named Entity Recognition (DNER) as a new task. It proposes Joint Argument and Entity Tagging (JAET), a generative framework that fine‑tunes decoder‑only LLMs to insert inline argument and entity tags into debate turns while preserving the original transcript. JAET achieves significant improvements in joint AM+DNER performance (+27.3% relative F1 in the untyped setting and +41.9% in the typed setting) over sequential pipelines, and these gains generalize to Persuasive Essays (+26.6% and +52.7%).

By Lucio La Cava, Stefano Francesco Monea, Sergio Greco
Hugging Face Trending Papers
Jun 10

A Resource for Enthymeme Detection in Controversial Political Discourse

Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjective. We present a resource of 1,482 tweets from politically controversial discourse, annotated by five annotators for the presence of enthymemes and their argument structure, designed to study label variation.

arXiv AI
Jun 17

RooseBERT: A New Deal For Political Language Modelling

arXiv:2508. 03250v4 Announce Type: replace-cross Abstract: The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to automatically analyse such content with the final goal of lightening up political deliberation to citizens.

By Deborah Dore, Elena Cabrio, Serena Villata
arXiv Computation and Language
Sep 14

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.

By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv Computation and Language
Aug 31

Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis

The paper introduces a retrieval‑augmented framework for detecting and classifying fallacies in political debate transcripts. By dynamically retrieving documents guided by argumentative relations of support and attack, the method leverages external knowledge to improve performance. Experiments on the ElecDeb60to20 benchmark show significant gains, raising macro‑F1 to 0.864 for detection and 0.725 for classification compared to non‑retrieval baselines.

By Deborah Dore, Greta Damo, Elena Cabrio, Serena Villata
arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin