arXiv AI

DSL-Topic: Improving Topic Modeling by Distilling Soft Labelsfrom Language Models

arXiv:2602. 17907v2 Announce Type: replace-cross Abstract: Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with data sparsity.

arXiv AI
Sep 1

Label Semantic Expansion via Label Guided Neural Topic Modeling

The paper introduces Label Semantic Expansion (LSE), a method that enriches sparse label representations by adding descriptive topic words grounded in a corpus. It proposes a Label-Guided Neural Topic Model (LGNTM) that learns label-aligned topics, integrates lexical and document semantics, and maintains consistency between topic and label structures. Experiments show that LSE and LGNTM improve label-topic alignment, label expansion, topic quality, and downstream classification performance.

By Haojia Zheng, Yuyin Lu, Juntian Huang, Fan Ou, Yanghui Rao, Haoran Xie, Fu Lee Wang
arXiv Computation and Language
Sep 10

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.

By Una Joh, Bei Yu
arXiv Computation and Language
Aug 25

Dynamic Topic Modeling for Cross-Corpus Temporal Analysis

Dynamic Embedded Topic Models (D-ETM) are extended to enable stable cross‑corpus temporal analysis by first learning a shared dynamic topic space—called the shared backbone—over a merged multi‑corpus collection. Corpus‑specific residual adaptation is then applied around this frozen backbone, allowing each corpus to specialize lexically without creating separate latent topic spaces. Experiments on three corpora spanning 97 years show that this approach yields much stronger alignment of topic trajectories (97.5 ± 0.7 % Retrieval@1) compared to full fine‑tuning or independent training with post‑hoc matching.

By Ruoxuan Li, Bruce Kogut
arXiv AI
6d ago

Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency

The paper introduces SeLATM, a framework that improves topic modeling by generating topics at the segment level and refining them through agentic feedback loops. This approach addresses limitations of LLM-based topic assignment methods, such as the inability to produce topic distributions, overly broad or narrow topics, and high resource consumption. Experiments on multiple datasets show that SeLATM reduces LLM resource usage while maintaining superior performance.

By Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong
arXiv AI
Aug 28

Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation

The paper examines how new vocabulary tokens are added to language models for generative recommendation tasks. It shows that the common practice of initializing these tokens as the mean of existing embeddings collapses them into a degenerate subspace, hindering fine‑tuning. The authors propose Grounded Token Initialization (GTI), which places new tokens at semantically meaningful positions in the pretrained embedding space using linguistic supervision, and demonstrate that GTI outperforms mean initialization and other adaptation methods across several benchmarks.

By Daiwei Chen, Zhoutong Fu, Chengming Jiang, Haichao Zhang, Ran Zhou, Tan Wang, Chunnan Yao, Guoyao Li, Rui Cai, Yihan Cao, Ruijie Jiang, Fedor Borisyuk, Jianqiang Shen, Jingwei Wu, Ramya Korlakai Vinayak
arXiv AI
Aug 25

Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling

The paper introduces LLM-QL, a dense retrieval model that harnesses large language models (LLMs) by maximizing query likelihood (QL) as an auxiliary task. It incorporates an Attention Block to limit predictive token attention to document tokens before the ending token and a Document Corruption component that masks parts of the document during prediction. Experiments on MS MARCO and BEIR datasets show that LLM-QL outperforms other LLM-based retrievers, and detailed analyses confirm the effectiveness of its components.

By Hengran Zhang, Keping Bi, Jiafeng Guo, Xiaojie Sun, Shihao Liu, Daiting Shi, Dawei Yin, Xueqi Cheng
arXiv Machine Learning
Sep 1

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

The paper investigates why token prediction, a common pre‑training objective for language models, yields useful representations. It introduces a statistical framework linking token prediction accuracy to the geometry of token embeddings, showing that accurate predictions organize embeddings according to Hellinger distances between context distributions. The authors also propose a self‑consistency principle that refines contextual representations through repeated application of a shared block, and provide downstream guarantees for token generation, community recovery, and linear classification.

By Shulei Wang
arXiv Machine Learning
Sep 10

Retrieval-augmented Decoding for Improving Truthfulness in Open-ended Generation

The paper introduces Retrieval-Augmented Decoding (RAD), a decoding-time method that improves the truthfulness of large language models without retraining. RAD uses a small reference set of up to ten annotated examples to build a grounding space of context embeddings and next-token logits, which it retrieves and aggregates during inference to shape the model’s output. Experiments on four open-ended generation benchmarks and four different LLMs show that RAD consistently outperforms strong baselines and generalizes well across tasks.

By Manh Nguyen, Sunil Gupta, Hung Le