Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
arXiv:2507. 23220v2 Announce Type: replace-cross Abstract: Traditional topic models are effective at uncovering latent themes in large text collections.
The paper introduces SeLATM, a framework that improves topic modeling by generating topics at the segment level and refining them through agentic feedback loops. This approach addresses limitations of LLM-based topic assignment methods, such as the inability to produce topic distributions, overly broad or narrow topics, and high resource consumption. Experiments on multiple datasets show that SeLATM reduces LLM resource usage while maintaining superior performance.
arXiv:2507. 23220v2 Announce Type: replace-cross Abstract: Traditional topic models are effective at uncovering latent themes in large text collections.
arXiv:2602. 17907v2 Announce Type: replace-cross Abstract: Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity.
arXiv:2602. 17907v3 Announce Type: replace-cross Abstract: Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity.
arXiv:2510. 16152v2 Announce Type: replace-cross Abstract: Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science.
TopiCLEAR is a framework that clusters document or sentence embeddings using adaptive dimensionality reduction to uncover low‑dimensional geometric structures that correspond to human‑interpretable topics. The method is evaluated on four benchmark datasets, showing strong agreement with human annotations, especially for short and informal texts. A Twitter case study demonstrates that TopiCLEAR yields more interpretable topics than LDA, recovering both annotated topic structure and coherent sub‑topics.
arXiv:2510. 18908v2 Announce Type: replace-cross Abstract: Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID-19 pandemic.
arXiv:2603.18358v2 Announce Type: replace Abstract: Outliers in dynamic topic modeling are typically treated as noise, yet we show that some can serve as early signals of emerging topics. We introduc...
The paper introduces MARETopic, a training‑free framework that identifies topics by selecting rank‑based prototype documents from pretrained embeddings. By projecting embeddings onto a low‑dimensional manifold and building ranked neighborhood lists, a greedy algorithm picks exactly K exemplar texts whose neighborhoods cover the corpus. Two variants—MARETopic_Corr, which uses a query‑performance predictor and rank correlation, and MARETopic_Diff, which employs a rank‑based diffusion matrix—achieve higher purity and NMI on benchmark datasets and run significantly faster, while also improving topic coherence and vocabulary diversity through a novel Maximal Marginal Relevance step.
The paper introduces Label Semantic Expansion (LSE), a method that enriches sparse label representations by adding descriptive topic words grounded in a corpus. It proposes a Label-Guided Neural Topic Model (LGNTM) that learns label-aligned topics, integrates lexical and document semantics, and maintains consistency between topic and label structures. Experiments show that LSE and LGNTM improve label-topic alignment, label expansion, topic quality, and downstream classification performance.
arXiv:2606. 27394v1 Announce Type: cross Abstract: The exponential increase in scientific publications has driven the emergence of new trends.
arXiv:2604.01452v2 Announce Type: replace Abstract: Scientific discovery is slowed by fragmented literature that requires excessive human effort to gather, analyze, and understand. AI tools, includin...
arXiv:2606. 10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions.