arXiv:2602. 17907v3 Announce Type: replace-cross Abstract: Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity.
By Raymond Li, Amirhossein Abaskohi, Chuyuan Li, Gabriel Murray, Giuseppe Carenini
The paper introduces Label Semantic Expansion (LSE), a method that enriches sparse label representations by adding descriptive topic words grounded in a corpus. It proposes a Label-Guided Neural Topic Model (LGNTM) that learns label-aligned topics, integrates lexical and document semantics, and maintains consistency between topic and label structures. Experiments show that LSE and LGNTM improve label-topic alignment, label expansion, topic quality, and downstream classification performance.
By Haojia Zheng, Yuyin Lu, Juntian Huang, Fan Ou, Yanghui Rao, Haoran Xie, Fu Lee Wang
arXiv:2608. 16269v1 Announce Type: cross Abstract: Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora.
By Seung-Won Seo, Won Ik Cho, Yongmin Yoo
arXiv:2507. 23220v2 Announce Type: replace-cross Abstract: Traditional topic models are effective at uncovering latent themes in large text collections.
By Carolina Zheng, Nicolas Beltran-Velez, Sweta Karlekar, Claudia Shi, Achille Nazaret, Asif Mallik, Amir Feder, David M. Blei
MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.
By Una Joh, Bei Yu
Dynamic Embedded Topic Models (D-ETM) are extended to enable stable cross‑corpus temporal analysis by first learning a shared dynamic topic space—called the shared backbone—over a merged multi‑corpus collection. Corpus‑specific residual adaptation is then applied around this frozen backbone, allowing each corpus to specialize lexically without creating separate latent topic spaces. Experiments on three corpora spanning 97 years show that this approach yields much stronger alignment of topic trajectories (97.5 ± 0.7 % Retrieval@1) compared to full fine‑tuning or independent training with post‑hoc matching.
By Ruoxuan Li, Bruce Kogut