arXiv Machine Learning

TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies

arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.

arXiv Computation and Language
Aug 28

TopiCLEAR: Adaptive embedding clustering for interpretable topic discovery from short texts

TopiCLEAR is a framework that clusters document or sentence embeddings using adaptive dimensionality reduction to uncover low‑dimensional geometric structures that correspond to human‑interpretable topics. The method is evaluated on four benchmark datasets, showing strong agreement with human annotations, especially for short and informal texts. A Twitter case study demonstrates that TopiCLEAR yields more interpretable topics than LDA, recovering both annotated topic structure and coherent sub‑topics.

By Aoi Fujita, Taichi Yamamoto, Yuri Nakayama, Ryota Kobayashi
arXiv Computation and Language
Sep 22

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.

By Simon Malan, Danel Slabbert, Herman Kamper
arXiv Machine Learning
Sep 7

Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale

The paper presents a scalable two‑stage clustering algorithm that guarantees per‑sample quality guardrails for large‑scale LLM‑based recommender systems. By first forming Mini‑batch K‑Means clusters and then greedily selecting representatives that meet user‑specified similarity and attribute constraints, the method ensures each sample inherits only relevant and safe outputs. Benchmarks show the approach runs faster and scales to millions of inputs, achieving a 50‑fold reduction in downstream LLM cost and runtime while maintaining personalization.

By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer
arXiv Computation and Language
Sep 15

Recovering the Zipfian Distribution in Unsupervised Term Discovery

The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.

By Danel Slabbert, Simon Malan, Herman Kamper
arXiv AI
Sep 25

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

The paper introduces MARETopic, a training‑free framework that identifies topics by selecting rank‑based prototype documents from pretrained embeddings. By projecting embeddings onto a low‑dimensional manifold and building ranked neighborhood lists, a greedy algorithm picks exactly K exemplar texts whose neighborhoods cover the corpus. Two variants—MARETopic_Corr, which uses a query‑performance predictor and rank correlation, and MARETopic_Diff, which employs a rank‑based diffusion matrix—achieve higher purity and NMI on benchmark datasets and run significantly faster, while also improving topic coherence and vocabulary diversity through a novel Maximal Marginal Relevance step.

By Thiago C\'esar Castilho Almeida, Daniel Carlos Guimar\~aes Pedronette