arXiv Computation and Language By Manuel Tonneau, Abhinav Dubey, Farhan Shaikh, Ilaria Vitulano, Martha Stolze, Hale Dedeoglu, Clara Riechert, Ella Kuka, Maryna Sydorova, Mykola Makhortykh, Elizaveta Kuznetsova

SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

Read the original on arXiv Computation and Language →

The paper introduces SWARM, a multilingual dataset of 2,183 search engine results in nine languages, annotated for support of Russian propaganda narratives. It evaluates a source-based blocklist, supervised classifiers, and zero‑shot large language models, finding that blocklists miss most propaganda and that content‑level models vary in performance, with the best LLM achieving an F1 of 0.73. The study highlights the need for per‑language, content‑level detection of search‑borne propaganda.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 25

ProBel: Propaganda Detection with Techniques, Spans, and Explanations

ProBel is a bilingual Arabic and English resource for propaganda detection that aligns binary labels, multi-label annotations for 23 propaganda techniques grouped into six categories, technique-labeled spans, and reference explanations for news sentences. The dataset supports matched binary, coarse-grained, multi-label, and span-level tasks in both languages, and the authors evaluate zero‑shot prompting, task‑specific fine‑tuning, and joint training. A single bilingual multi‑task model achieves the best overall performance, with cross‑task analysis revealing that joint classification preserves binary performance while span‑only training can weaken sentence‑level prediction, and that joint bilingual training yields the most stable results.

By Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori, Giovanni Da San Martino, Firoj Alam
arXiv AI
Aug 18

Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign

arXiv:2608. 15746v1 Announce Type: new Abstract: We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign.

By Benjamin Icard, Elouan Vuichard, Louis Lefebvre, Lila Sainero, Thomas Girault, Alice Breton, Tanguy Launay, Gauvain Bourgne, Morgane Casanova, Guillaume Gadek, Victor Kl\"otzer, Michel Le Nouy, Guillaume Gravier, Jean-Gabriel Ganascia, Paul \'Egr\'e
arXiv Computation and Language
2d ago

Zero-shot narrative detection in social messaging

The paper explores how large language models can detect hidden narratives in social messages without training data. By feeding the models human-written narrative descriptions, performance improves markedly, while automatically generated descriptions or few-shot examples can hurt accuracy. Ensemble techniques, especially majority voting, further boost robustness, and larger models show the best results with less sensitivity to prompts.

By Jes\'us M. Fraile-Hern\'andez, Anselmo Pe\~nas, Patrick Giedemann
arXiv Machine Learning
Aug 24

MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees

MIL-BERT is a neural network algorithm that classifies large texts by selecting relevant excerpts, inspired by multiple instance learning. It scales to samples with nearly 1 million tokens and has been evaluated on seven datasets, achieving state‑of‑the‑art results on three long‑text tasks such as political bias detection, trigger warning identification, and author demographic inference. The model also generalizes from weakly‑labeled text bags to accurately classify smaller instances.

By John Cadigan, Dayne Freitag, Eric Yeh
arXiv Computation and Language
Sep 7

Multilingual Models for Check-Worthy Social Media Posts Detection

The paper reports a comprehensive study of transformer-based NLP models for detecting check-worthy social media posts, covering data collection, preprocessing, architecture selection, fine‑tuning, testing, and implementation. It focuses on multilingual models that can process English and low‑resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, and Slovak, and compares their performance to state‑of‑the‑art baselines. The work introduces multi‑label multilingual classifiers that simultaneously identify harmful content and posts containing verifiable factual claims efficiently.

By Sebastian Kula