arXiv Machine Learning

Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate Discourse

arXiv:2601. 13317v2 Announce Type: replace-cross Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally different environments: paid advertising platforms host targeted, institutionally produced messaging, while public social media reflects largely organic, user-driven discussion.

arXiv Computation and Language
Sep 14

Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework

The paper introduces an open, modular AI framework that automatically detects and structures evidence of social tipping points in climate literature at the passage level. It integrates a DistilBERT boundary splitter, an iteratively augmented RoBERTa classifier, a Mistral 7B rewrite model, a LLaMA 3.2 3B rating model, and a Milvus vector store, all accessible via a Streamlit interface. Evaluation on a GPT‑4.1‑labelled benchmark and expert‑reviewed set shows the splitter outperforms competitors and the RoBERTa detector achieves high accuracy and agreement.

By Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicol\`o Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian
arXiv Computation and Language
Sep 17

Structured Claim-Level Discourse Representations for Dense Health Narratives

The paper introduces a structured framework for claim-level discourse analysis in dense health narratives, addressing the limitations of existing topic- or sentiment-based representations. It identifies an average of 13.22 atomic claims per minute in social media health videos and proposes tuples that link each claim to thematic aspects, stance, and multidimensional pragmatic attributes. A benchmark of 1,191 manually annotated claims from 60 videos across four health domains is created, and experiments show that large language models perform well on thematic categorization and stance prediction but struggle with high-dimensional pragmatic profiling, indicating a need for task-specific inference strategies.

By Farnoushsadat Nilizadeh, Elham Pourabbas Vafa, Shirin Nilizadeh, Eduard Dragut
arXiv Computation and Language
Sep 11

Automated Identification of Competing Narratives in Political Discourse on Social Media

The paper introduces an unsupervised framework that identifies and characterizes competing narratives in political discourse on social media, specifically analyzing German politicians' tweets. It uses a multi‑stage pipeline incorporating topic modeling, event detection, and event linking to form coherent stories and reveal distinct user community perspectives. Two case studies on polarizing issues demonstrate the method’s effectiveness in uncovering divergent viewpoints and framing conflicts around trending political topics.

By Sergej Wildemann, Erick Elejalde
arXiv Machine Learning
Sep 25

Agentic Detection of Online Conspiracies

The paper presents an agentic framework for detecting conspiratorial content in social media by inferring the speaker’s intent rather than merely identifying explicit claims. It leverages social context and adaptive tool use, demonstrating superior performance over text-only and non-agentic models on a large Hebrew tweet dataset spanning election cycles and the COVID pandemic. The study highlights the importance of context-aware, reasoning-driven approaches for accurate conspiracy detection.

By Lior Biton, Oren Tsur
arXiv AI
Jul 21

Posts of Peril: Detecting Information About Hazards in Text

arXiv:2405. 17838v3 Announce Type: replace-cross Abstract: Socio-linguistic indicators of affectively-relevant phenomena, such as emotion or sentiment, are often extracted from text to better understand features of human-computer interactions, including on social media.

By Keith Burghardt, Daniel M. T. Fessler, Chyna Tang, Anne Pisor, Kristina Lerman
arXiv AI
Aug 7

Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation

arXiv:2608. 05155v1 Announce Type: cross Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to research in the social sciences and humanities (SSH).

By Maryam Fooladi, Federico Bottino
arXiv Machine Learning
Aug 20

BERTilda: Explainable Topic Lifecycle Tracking with Split/Merge Detection via Similarity-and-Flow Temporal Graphs

BERTilda is an explainable framework for tracking topic lifecycles in longitudinal text streams. It discovers topics independently in each time window using an embedding‑based topic model, then links topics across adjacent windows via a temporal graph that uses both semantic similarity and a bidirectional coverage signal derived from tweet‑to‑topic attribution. The graph‑based rules identify continuations, splits, merges, disappearances, and unclear transitions, and the method achieves up to 87% agreement with human annotators on a gold‑standard subset.

By Cl\'audia Oliveira, \'Alvaro Figueira