arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.
By Daoming Wan, Yizheng Huang, Jimmy X. Huang
arXiv:2506. 10292v2 Announce Type: replace-cross Abstract: Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance on extensive labelled data.
By Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi, Imran Razzak
The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.
By Simon Malan, Danel Slabbert, Herman Kamper
arXiv:2607. 28635v1 Announce Type: cross Abstract: In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics.
By Noor Khalal, Abdallah Alaa-Eddine Djamai, Imed Keraghel, Mohamed Nadif
arXiv:2607. 17653v1 Announce Type: cross Abstract: Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data.
By Jing Li, Pan Liu, Meng Zhao, Wanli Xue, Yanhong Yang, Xu Cheng, Fan Shi, Jianhua Zhang, Qinghua Hu, Shengyong Chen
arXiv:2609.24464v1 Announce Type: new
Abstract: Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial pr...
By Armin Oliya, Aleksandra Sawczuk, Rados{\l}aw Bia{\l}obrzeski
Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real wor...
arXiv:2607. 21426v1 Announce Type: new Abstract: Cooperative multi-task semantic communication (CMT-SemCom) improves task execution performance by leveraging shared representations.
By Ahmad Halimi Razlighi, Maximilian H. V. Tillmann, Edgar Beck, Bho Matthiesen, Armin Dekorsy
arXiv:2602. 11641v2 Announce Type: replace Abstract: Text-attributed graphs (TAGs) associate nodes with textual attributes and graph structure, enabling GNNs to jointly model semantic and structural information.
By Yinlin Zhu, Di Wu, Xu Wang, Guocong Quan, Miao Hu
arXiv:2607. 02266v1 Announce Type: cross Abstract: Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer can express.
By Ziyun Qiao, Yue Min, Ruining Chen, Yujun Li
The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.
By Danel Slabbert, Simon Malan, Herman Kamper
arXiv:2511. 05913v2 Announce Type: replace-cross Abstract: New intent discovery (NID) seeks to recognize both new and known intents from unlabeled user utterances, which finds prevalent use in practical dialogue systems.
By Hongtao Wang, Renchi Yang, Wenqing Lin