arXiv Machine Learning By Parth Bramhecha, Smit Deshmukh, Sairaj Bodhale, Adwait Borate, Raviraj Joshi

BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events

Read the original on arXiv Machine Learning →

BharatGather is a curated, multi-source dataset designed for binary misinformation classification in Indian public events such as religious festivals, political rallies, and cultural gatherings. The corpus contains 14,646 records assembled through systematic web scraping of fact‑checking platforms, multimedia transcript extraction, and LLM‑mediated synthetic augmentation to capture narrative diversity. It serves as a culturally informed benchmark to evaluate and develop fake‑news detection systems tailored to the socio‑cultural nuances of India’s mass‑gathering context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
2d ago

FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool

FakeSpotter is a new tool that estimates the viral misinformation risk of textual content by measuring structural fingerprints of misinformation instead of directly judging truthfulness. It operates across linguistic, narrative, logical, and critical‑thinking dimensions, using repeated large language model assessments and domain‑specific logistic regression classifiers for both short and long texts. In a labeled corpus of 764 texts, FakeSpotter achieved macro F1 scores of 0.788 for short texts and 0.793 for long texts, and its interpretive layer offers explainable outputs such as feature‑based scores, signal agreement, and a caution index for social listening.

By Giovanni Spitale, Federico Germani
arXiv AI
Aug 25

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

The paper introduces BanglaSafe, a benchmark of 879 Bengali prompts that covers 17 culturally grounded harm categories and five prompting conditions. Evaluation of 18 frontier LLMs shows that 53.6% of responses are unsafe or partially unsafe, with 14.7% containing strictly harmful content. The study finds that the writing style within Bengali has a stronger impact on safety than the language switch itself, and that current safety classifiers struggle to reliably evaluate Bengali content.

By Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady
Hugging Face Trending Papers
Sep 2

From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X

The paper presents a large‑scale study of ragebait—content designed to provoke anger—on Japanese posts on X. It introduces a labeled dataset created with a large language model, trains Japanese language models, and builds an ensemble classifier that detects ragebait. Applying this detector to a vast dataset reveals that ragebait is especially common in politically and socially contentious topics, spreads faster, and elicits stronger negative emotions than non‑ragebait posts.

arXiv Computation and Language
Sep 3

From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X

The paper presents a large‑scale study of ragebait on Japanese X, developing an ensemble classifier trained on a dataset labeled with the help of a large language model. The detector was applied to a vast collection of Japanese posts, revealing that ragebait is especially common in politically and socially contentious topics such as politics, discrimination, public health, and interpersonal conflict. Ragebait posts spread more quickly and elicit stronger negative emotions—anger, fear, disgust, sadness, and surprise—than non‑ragebait posts.

By Zhiyang Qi, Kazuhiro Ito, Jinghui Chen, Hibiki Nakamura, Zhangxuan Chen, Erina Murata, Masaki Chujyo, Fujio Toriumi