ProBel is a bilingual Arabic and English resource for propaganda detection that aligns binary labels, multi-label annotations for 23 propaganda techniques grouped into six categories, technique-labeled spans, and reference explanations for news sentences. The dataset supports matched binary, coarse-grained, multi-label, and span-level tasks in both languages, and the authors evaluate zero‑shot prompting, task‑specific fine‑tuning, and joint training. A single bilingual multi‑task model achieves the best overall performance, with cross‑task analysis revealing that joint classification preserves binary performance while span‑only training can weaken sentence‑level prediction, and that joint bilingual training yields the most stable results.
By Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Elisa Sartori, Giovanni Da San Martino, Firoj Alam
arXiv:2608. 15746v1 Announce Type: new Abstract: We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign.
By Benjamin Icard, Elouan Vuichard, Louis Lefebvre, Lila Sainero, Thomas Girault, Alice Breton, Tanguy Launay, Gauvain Bourgne, Morgane Casanova, Guillaume Gadek, Victor Kl\"otzer, Michel Le Nouy, Guillaume Gravier, Jean-Gabriel Ganascia, Paul \'Egr\'e
The paper explores how large language models can detect hidden narratives in social messages without training data. By feeding the models human-written narrative descriptions, performance improves markedly, while automatically generated descriptions or few-shot examples can hurt accuracy. Ensemble techniques, especially majority voting, further boost robustness, and larger models show the best results with less sensitivity to prompts.
By Jes\'us M. Fraile-Hern\'andez, Anselmo Pe\~nas, Patrick Giedemann
MIL-BERT is a neural network algorithm that classifies large texts by selecting relevant excerpts, inspired by multiple instance learning. It scales to samples with nearly 1 million tokens and has been evaluated on seven datasets, achieving state‑of‑the‑art results on three long‑text tasks such as political bias detection, trigger warning identification, and author demographic inference. The model also generalizes from weakly‑labeled text bags to accurately classify smaller instances.
By John Cadigan, Dayne Freitag, Eric Yeh
The paper reports a comprehensive study of transformer-based NLP models for detecting check-worthy social media posts, covering data collection, preprocessing, architecture selection, fine‑tuning, testing, and implementation. It focuses on multilingual models that can process English and low‑resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, and Slovak, and compares their performance to state‑of‑the‑art baselines. The work introduces multi‑label multilingual classifiers that simultaneously identify harmful content and posts containing verifiable factual claims efficiently.
By Sebastian Kula
arXiv:2607. 20463v1 Announce Type: new Abstract: This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles.
By Wojciech Michaluk, Tymoteusz Urban, Mateusz Kubita, Soveatin Kuntur, Anna Wr\'oblewska
The study analyzes millions of German-language online articles and tweets from 2019–2022 to uncover political biases using automated text analysis. It finds that international events such as the COVID‑19 pandemic and the Ukraine war create thematic convergence between German and Swiss media, while domestic policy differences drive divergence in locally focused topics. Newspapers maintain more stable political content, whereas Twitter shows rapid, event‑driven spikes, illustrating how media platforms differ in intensity and timing.
By Yara D\"oring, Felix Bie{\ss}mann
The paper presents a large‑scale study of ragebait—content designed to provoke anger—on Japanese posts on X. It introduces a labeled dataset created with a large language model, trains Japanese language models, and builds an ensemble classifier that detects ragebait. Applying this detector to a vast dataset reveals that ragebait is especially common in politically and socially contentious topics, spreads faster, and elicits stronger negative emotions than non‑ragebait posts.
arXiv:2608.22922v1 Announce Type: new
Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourc...
By Thisen Ekanayake, Nisansa de Silva
The paper presents a large‑scale study of ragebait on Japanese X, developing an ensemble classifier trained on a dataset labeled with the help of a large language model. The detector was applied to a vast collection of Japanese posts, revealing that ragebait is especially common in politically and socially contentious topics such as politics, discrimination, public health, and interpersonal conflict. Ragebait posts spread more quickly and elicit stronger negative emotions—anger, fear, disgust, sadness, and surprise—than non‑ragebait posts.
By Zhiyang Qi, Kazuhiro Ito, Jinghui Chen, Hibiki Nakamura, Zhangxuan Chen, Erina Murata, Masaki Chujyo, Fujio Toriumi
arXiv:2608. 03859v1 Announce Type: cross Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review.
By Peijia Guo, Wenxuan Xie, ZiGuang Li, Ming Li
The study examines how duplicate content is used to detect coordination in social media information operations. It distinguishes between generic, low‑information duplicates and non‑generic, more specific duplicates, labeling 187,000 tweets with an LLM‑assisted protocol and supervised classifiers. Results show that generic duplicates are rare with lexical matching but comprise nearly 39% of campaigns identified by embedding methods, and filtering out generic duplicates yields smaller, denser coordination graphs, indicating a more focused structure.
By Ashfaq Ali Shafin, Khandaker Mamun Ahmed