arXiv:2603.18358v2 Announce Type: replace
Abstract: Outliers in dynamic topic modeling are typically treated as noise, yet we show that some can serve as early signals of emerging topics. We introduc...
By Evangelia Zve, Gauvain Bourgne, Benjamin Icard, Jean-Gabriel Ganascia
The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.
By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia
The paper introduces a new method that uses embedding-based dynamic topic modeling to detect and measure topic drift at the comment level in a large dataset. By applying pretrained language models to generate contextualized embeddings for 12.7 billion Reddit comments from 2006 to 2022, the authors identify evolving topic clusters over time using unsupervised techniques. Their approach includes scalable modifications to existing methods and a null model comparison test, revealing that politically and socially contentious topics show significant directional drift, while areas like music and sports remain relatively stable.
The paper presents an end‑to‑end framework for extracting and clustering trilingual Sri Lankan parliamentary debates in Sinhala, Tamil, and English. Using LLM‑based text extraction, multilingual embeddings, and density‑based clustering, the authors recover 30 macro‑topics with a cluster purity of 0.673. The temporal patterns of these topics align with major national events such as the 2019 Easter attacks and the 2022 economic crisis, demonstrating the method’s effectiveness where traditional LDA fails.
By Himath Dhanapala, Haren Daishika, Himandhi Kuruppu, Sithija Seneviratne, Ashini Kavindya, Patalee Narasinghe, Sandeepa Weerasekara, Nisansa de Silva, Sandareka Wickramanayake
BERTilda is an explainable framework for tracking topic lifecycles in longitudinal text streams. It discovers topics independently in each time window using an embedding‑based topic model, then links topics across adjacent windows via a temporal graph that uses both semantic similarity and a bidirectional coverage signal derived from tweet‑to‑topic attribution. The graph‑based rules identify continuations, splits, merges, disappearances, and unclear transitions, and the method achieves up to 87% agreement with human annotators on a gold‑standard subset.
By Cl\'audia Oliveira, \'Alvaro Figueira
arXiv:2510. 18908v2 Announce Type: replace-cross Abstract: Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic.
By Wangjiaxuan Xin, Shuhua Yin, Shi Chen, Yaorong Ge
arXiv:2608. 06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects.
By Karolina Rudnicka, Thomas Stephan Juzek
The study analyzes millions of German-language online articles and tweets from 2019–2022 to uncover political biases using automated text analysis. It finds that international events such as the COVID‑19 pandemic and the Ukraine war create thematic convergence between German and Swiss media, while domestic policy differences drive divergence in locally focused topics. Newspapers maintain more stable political content, whereas Twitter shows rapid, event‑driven spikes, illustrating how media platforms differ in intensity and timing.
By Yara D\"oring, Felix Bie{\ss}mann
MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.
By Una Joh, Bei Yu
arXiv:2601. 13317v2 Announce Type: replace-cross Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally different environments: paid advertising platforms host targeted, institutionally produced messaging, while public social media reflects largely organic, user-driven discussion.
By Samantha Sudhoff, Pranav Perumal, Zhaoqing Wu, Tunazzina Islam
Dynamic Embedded Topic Models (D-ETM) are extended to enable stable cross‑corpus temporal analysis by first learning a shared dynamic topic space—called the shared backbone—over a merged multi‑corpus collection. Corpus‑specific residual adaptation is then applied around this frozen backbone, allowing each corpus to specialize lexically without creating separate latent topic spaces. Experiments on three corpora spanning 97 years show that this approach yields much stronger alignment of topic trajectories (97.5 ± 0.7 % Retrieval@1) compared to full fine‑tuning or independent training with post‑hoc matching.
By Ruoxuan Li, Bruce Kogut
arXiv:2609.08609v2 Announce Type: replace
Abstract: Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and t...
By Nevidu Jayatilleke, Nisansa de Silva