arXiv AI

ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models

arXiv:2607. 20463v1 Announce Type: new Abstract: This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles.

arXiv Machine Learning
Jun 16

YTClickbait21K: Human-Annotated Multimodal Dataset for YouTube Clickbait Detection Across Diverse Channels and Content Categories

arXiv:2606. 14780v1 Announce Type: cross Abstract: Clickbait content on video-sharing platforms poses a significant challenge to information reliability, yet progress in automated detection has been constrained by the lack of large-scale, high-quality multimodal datasets.

By Md. Minhazul Islam, Md. Tanbeer Jubaer, Amith Khandakar, Shovon Sarker, Sumaiya Rahman, Md. Masum Mia, Mohamed Arselene Ayari, Hamed Noori
Hugging Face Trending Papers
Sep 2

From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X

The paper presents a large‑scale study of ragebait—content designed to provoke anger—on Japanese posts on X. It introduces a labeled dataset created with a large language model, trains Japanese language models, and builds an ensemble classifier that detects ragebait. Applying this detector to a vast dataset reveals that ragebait is especially common in politically and socially contentious topics, spreads faster, and elicits stronger negative emotions than non‑ragebait posts.

arXiv Computation and Language
Sep 3

From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X

The paper presents a large‑scale study of ragebait on Japanese X, developing an ensemble classifier trained on a dataset labeled with the help of a large language model. The detector was applied to a vast collection of Japanese posts, revealing that ragebait is especially common in politically and socially contentious topics such as politics, discrimination, public health, and interpersonal conflict. Ragebait posts spread more quickly and elicit stronger negative emotions—anger, fear, disgust, sadness, and surprise—than non‑ragebait posts.

By Zhiyang Qi, Kazuhiro Ito, Jinghui Chen, Hibiki Nakamura, Zhangxuan Chen, Erina Murata, Masaki Chujyo, Fujio Toriumi
arXiv Computation and Language
Sep 14

SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

The paper introduces SWARM, a multilingual dataset of 2,183 search engine results in nine languages, annotated for support of Russian propaganda narratives. It evaluates a source-based blocklist, supervised classifiers, and zero‑shot large language models, finding that blocklists miss most propaganda and that content‑level models vary in performance, with the best LLM achieving an F1 of 0.73. The study highlights the need for per‑language, content‑level detection of search‑borne propaganda.

By Manuel Tonneau, Abhinav Dubey, Farhan Shaikh, Ilaria Vitulano, Martha Stolze, Hale Dedeoglu, Clara Riechert, Ella Kuka, Maryna Sydorova, Mykola Makhortykh, Elizaveta Kuznetsova
arXiv Machine Learning
5d ago

Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation

The paper presents a theory-informed computational framework that converts cross-disciplinary theories of fake news into measurable features for automated detection and explanation. By reviewing theories from social sciences, psychology, economics, and more, the authors establish a broad theoretical foundation for computational modeling. Experiments on benchmark datasets demonstrate that theory-derived features are predictive, provide interpretable diagnostic signals, and that multi-feature models generally outperform individual features, though gains are modest.

By Zhaoyang Cao, Miriam Metzger, Reza Zafarani
arXiv Machine Learning
Aug 27

BanglaMamba: Exploring State Space Models for Bangla Fake News Detection

BanglaMamba explores Mamba-based State Space Models (SSMs) as a computationally efficient alternative for Bangla fake news detection. Compared to BanglaBERT and a custom BERT trained from scratch, BanglaMamba achieves a Macro‑F1 score of 0.9029, close to the 0.9057 of the custom BERT, while delivering 2.2× higher inference throughput and 49% lower peak GPU memory usage. Cross‑dataset evaluation shows BanglaBERT generalizes better, underscoring the value of large‑scale pretraining.

By M. K. Khalidi Siam
arXiv AI
4d ago

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

The paper investigates whether multimodal large language models (MLLMs) can generate and detect realistic multimodal fake news on social media. Using a multi‑agent framework—comprising a story agent, an image agent, and a critic agent—the authors produced over 9,000 paired multimodal news posts across science, health, and entertainment domains. They benchmarked 16 open‑ and closed‑source MLLMs for automated detection and found that most models fall far short of human accuracy, especially in identifying image authenticity, highlighting the need for stronger defenses against social media fake news.

By Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, Xiuzhen Zhang
arXiv AI
6d ago

A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models

The article surveys fake review detection research, focusing on how pre‑trained language models (PLMs) and large language models (LLMs) influence both the generation of deceptive reviews and their detection. It reviews 211 studies from 2018 to early 2026, categorizing methods by evidence source—such as review text, sentiment, rating behavior, temporal metadata, user‑product graphs, multimodal content, external knowledge, and LLM‑generated signals—and by fusion level. The survey traces the evolution from traditional machine learning to PLM‑based and LLM‑based approaches, evaluates performance on Amazon, Yelp, and OpSpam benchmarks, and highlights open challenges including adversarial generation, cross‑domain transfer, uncertainty‑aware fusion, robustness to missing sources, interpretability, and trustworthy evaluation of AI‑generated deceptive content.

By Fanji Yang (Guizhou University of Finance and Economics), Huiyao Chen (Harbin Institute of Technology), Xi Yu (Guizhou University of Finance and Economics), Meishan Zhang (Harbin Institute of Technology), Xiaohong Xiao (Guizhou University of Commerce), Mingsen Deng (Guizhou University of Finance and Economics)