arXiv Computation and Language

From Outliers to Topics in Language Models: Anticipating Trends in News Corpora

arXiv Computation and Language
5d ago

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.

By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia
Hugging Face Trending Papers
Aug 19

Comment-level Topic Drift Analysis in the Reddit Corpus

The paper introduces a new method that uses embedding-based dynamic topic modeling to detect and measure topic drift at the comment level in a large dataset. By applying pretrained language models to generate contextualized embeddings for 12.7 billion Reddit comments from 2006 to 2022, the authors identify evolving topic clusters over time using unsupervised techniques. Their approach includes scalable modifications to existing methods and a null model comparison test, revealing that politically and socially contentious topics show significant directional drift, while areas like music and sports remain relatively stable.

arXiv AI
Aug 24

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

The paper presents an end‑to‑end framework for extracting and clustering trilingual Sri Lankan parliamentary debates in Sinhala, Tamil, and English. Using LLM‑based text extraction, multilingual embeddings, and density‑based clustering, the authors recover 30 macro‑topics with a cluster purity of 0.673. The temporal patterns of these topics align with major national events such as the 2019 Easter attacks and the 2022 economic crisis, demonstrating the method’s effectiveness where traditional LDA fails.

By Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake
arXiv Machine Learning
Aug 20

BERTilda: Explainable Topic Lifecycle Tracking with Split/Merge Detection via Similarity-and-Flow Temporal Graphs

BERTilda is an explainable framework for tracking topic lifecycles in longitudinal text streams. It discovers topics independently in each time window using an embedding‑based topic model, then links topics across adjacent windows via a temporal graph that uses both semantic similarity and a bidirectional coverage signal derived from tweet‑to‑topic attribution. The graph‑based rules identify continuations, splits, merges, disappearances, and unclear transitions, and the method achieves up to 87% agreement with human annotators on a gold‑standard subset.

By Cl\'audia Oliveira, \'Alvaro Figueira
arXiv Machine Learning
Aug 20

Global Crises and National Policies: A Large Scale Analysis of Political Content in German Language Online Media

The study analyzes millions of German-language online articles and tweets from 2019–2022 to uncover political biases using automated text analysis. It finds that international events such as the COVID‑19 pandemic and the Ukraine war create thematic convergence between German and Swiss media, while domestic policy differences drive divergence in locally focused topics. Newspapers maintain more stable political content, whereas Twitter shows rapid, event‑driven spikes, illustrating how media platforms differ in intensity and timing.

By Yara D\"oring, Felix Bie{\ss}mann
arXiv Computation and Language
Sep 10

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.

By Una Joh, Bei Yu
arXiv Machine Learning
Jun 25

Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate Discourse

arXiv:2601. 13317v2 Announce Type: replace-cross Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally different environments: paid advertising platforms host targeted, institutionally produced messaging, while public social media reflects largely organic, user-driven discussion.

By Samantha Sudhoff, Pranav Perumal, Zhaoqing Wu, Tunazzina Islam
arXiv Computation and Language
Aug 25

Dynamic Topic Modeling for Cross-Corpus Temporal Analysis

Dynamic Embedded Topic Models (D-ETM) are extended to enable stable cross‑corpus temporal analysis by first learning a shared dynamic topic space—called the shared backbone—over a merged multi‑corpus collection. Corpus‑specific residual adaptation is then applied around this frozen backbone, allowing each corpus to specialize lexically without creating separate latent topic spaces. Experiments on three corpora spanning 97 years show that this approach yields much stronger alignment of topic trajectories (97.5 ± 0.7 % Retrieval@1) compared to full fine‑tuning or independent training with post‑hoc matching.

By Ruoxuan Li, Bruce Kogut