arXiv Machine Learning By Daoming Wan, Yizheng Huang, Jimmy X. Huang

TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies

Read the original on arXiv Machine Learning →

arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 28

TopiCLEAR: Adaptive embedding clustering for interpretable topic discovery from short texts

TopiCLEAR is a framework that clusters document or sentence embeddings using adaptive dimensionality reduction to uncover low‑dimensional geometric structures that correspond to human‑interpretable topics. The method is evaluated on four benchmark datasets, showing strong agreement with human annotations, especially for short and informal texts. A Twitter case study demonstrates that TopiCLEAR yields more interpretable topics than LDA, recovering both annotated topic structure and coherent sub‑topics.

By Aoi Fujita, Taichi Yamamoto, Yuri Nakayama, Ryota Kobayashi
arXiv Computation and Language
Sep 22

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.

By Simon Malan, Danel Slabbert, Herman Kamper