arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.
By Daoming Wan, Yizheng Huang, Jimmy X. Huang
arXiv:2511. 05913v2 Announce Type: replace-cross Abstract: New intent discovery (NID) seeks to recognize both new and known intents from unlabeled user utterances, which finds prevalent use in practical dialogue systems.
By Hongtao Wang, Renchi Yang, Wenqing Lin
arXiv:2608.30297v1 Announce Type: new
Abstract: Attributes describing data content and context can induce diverse imbalance patterns that go beyond label imbalance alone. However, existing studies pr...
By Hanshu Rao, Guangzeng Han, Xiaolei Huang
arXiv:2608. 00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy.
By Pulock Das, Yina Hou, Md. Kamrozzaman Bhuiyan, Manar D. Samad
LentEx is a new framework for latent entity extraction that uses synthetic data generation and instruction fine‑tuning to train smaller, efficient large language models. By creating diverse, contextually rich synthetic examples through a template‑based approach, LentEx overcomes the lack of labeled datasets and achieves strong performance, surpassing state‑of‑the‑art models on the MTEB Clustering Benchmark. The method also generalizes well to unseen domains, making it useful for tasks such as retrieval‑augmented generation, customer persona analysis, and knowledge graph enrichment.
By Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar, Shujing Dong, Yarong Feng, Hongfei Li, Ayush Goyal
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv:2607. 10312v1 Announce Type: cross Abstract: The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate effectively across diverse linguistic contexts.
By Muhammad Abdullahi Said
TopiCLEAR is a framework that clusters document or sentence embeddings using adaptive dimensionality reduction to uncover low‑dimensional geometric structures that correspond to human‑interpretable topics. The method is evaluated on four benchmark datasets, showing strong agreement with human annotations, especially for short and informal texts. A Twitter case study demonstrates that TopiCLEAR yields more interpretable topics than LDA, recovering both annotated topic structure and coherent sub‑topics.
By Aoi Fujita, Taichi Yamamoto, Yuri Nakayama, Ryota Kobayashi
MonoTM is an interpretable topic modeling framework that separates the estimation of document–topic mixtures from the generation of topic descriptors. It uses sparse autoencoders to extract dense, interpretable features for mixture estimation, then learns topic descriptors from a distinct set of corpus‑grounded semantic features. This approach preserves global topic structure while providing more meaningful, semantic‑unit descriptors than traditional top‑word lists.
By Una Joh, Bei Yu
arXiv:2507.12295v2 Announce Type: replace-cross
Abstract: Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation id...
By Feng Xiao, Jicong Fan
arXiv:2510. 09783v2 Announce Type: replace-cross Abstract: Oversampling is one of the most widely used approaches for addressing imbalanced classification.
By Dang Nguyen, Sunil Gupta, Kien Do, Thin Nguyen, Taylor Braund, Alexis Whitton, Svetha Venkatesh
We’ve obtained state-of-the-art results on a suite of diverse language tasks with a scalable, task-agnostic system, which we’re also releasing. Our approach is a combination of two existing ideas: transformers and unsupervised pre-training.