arXiv Computation and Language By Evangelia Zve, Gauvain Bourgne, Jean-Gabriel Ganascia

Predicting Emerging Topics from Outliers: A Prospective Study of Weak Signals in Embedding Space

Read the original on arXiv Computation and Language →

The study investigates whether documents initially classified as noise in embedding-based topic models can be identified as precursors to emerging topics. By labeling documents based on their future trajectories and measuring confidence across multiple embedding models, the authors find that anticipatory outliers are predictable at publication time, achieving an F1 score above 0.90 on high-consensus subsets and 0.76–0.80 in chronological evaluation. The predictive power largely stems from geometric features that capture each outlier’s position in embedding space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Aug 19

Comment-level Topic Drift Analysis in the Reddit Corpus

The paper introduces a new method that uses embedding-based dynamic topic modeling to detect and measure topic drift at the comment level in a large dataset. By applying pretrained language models to generate contextualized embeddings for 12.7 billion Reddit comments from 2006 to 2022, the authors identify evolving topic clusters over time using unsupervised techniques. Their approach includes scalable modifications to existing methods and a null model comparison test, revealing that politically and socially contentious topics show significant directional drift, while areas like music and sports remain relatively stable.

arXiv AI
5d ago

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

The paper introduces MARETopic, a training‑free framework that identifies topics by selecting rank‑based prototype documents from pretrained embeddings. By projecting embeddings onto a low‑dimensional manifold and building ranked neighborhood lists, a greedy algorithm picks exactly K exemplar texts whose neighborhoods cover the corpus. Two variants—MARETopic_Corr, which uses a query‑performance predictor and rank correlation, and MARETopic_Diff, which employs a rank‑based diffusion matrix—achieve higher purity and NMI on benchmark datasets and run significantly faster, while also improving topic coherence and vocabulary diversity through a novel Maximal Marginal Relevance step.

By Thiago C\'esar Castilho Almeida, Daniel Carlos Guimar\~aes Pedronette