arXiv Machine Learning By Ashfaq Ali Shafin, Khandaker Mamun Ahmed

Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations

Read the original on arXiv Machine Learning →

The study examines how duplicate content is used to detect coordination in social media information operations. It distinguishes between generic, low‑information duplicates and non‑generic, more specific duplicates, labeling 187,000 tweets with an LLM‑assisted protocol and supervised classifiers. Results show that generic duplicates are rare with lexical matching but comprise nearly 39% of campaigns identified by embedding methods, and filtering out generic duplicates yields smaller, denser coordination graphs, indicating a more focused structure.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
4d ago

SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

The paper introduces SWARM, a multilingual dataset of 2,183 search engine results in nine languages, annotated for support of Russian propaganda narratives. It evaluates a source-based blocklist, supervised classifiers, and zero‑shot large language models, finding that blocklists miss most propaganda and that content‑level models vary in performance, with the best LLM achieving an F1 of 0.73. The study highlights the need for per‑language, content‑level detection of search‑borne propaganda.

By Manuel Tonneau, Abhinav Dubey, Farhan Shaikh, Ilaria Vitulano, Martha Stolze, Hale Dedeoglu, Clara Riechert, Ella Kuka, Maryna Sydorova, Mykola Makhortykh, Elizaveta Kuznetsova
arXiv Machine Learning
Jun 11

GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs

arXiv:2606. 11898v1 Announce Type: cross Abstract: Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation networks, e-commerce platforms, social media, and web pages.

By Hengyi Feng, Zeang Sheng, Meiyi Qiang, Meiyi Qiang, Wentao Zhang
arXiv Computation and Language
Sep 11

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

The paper introduces a new perspective on entity rarity in multimodal entity linking by using knowledge‑graph structural metrics instead of popularity metrics, revealing many rare entities previously overlooked. Experiments show that state‑of‑the‑art models suffer a 15.4–39.9% accuracy drop on these rare‑entity slices. The authors propose a training‑free framework that combines reasoning and retrieval with a vision‑language model, achieving a 6.9% overall accuracy gain and up to 23.3% improvement on rare entities, and release a new benchmark MERLIN‑Rare for focused evaluation.

By Parinthapat Pengpun, Simran Khanuja, Graham Neubig