arXiv:2606. 28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems.
By Daoming Wan, Yizheng Huang, Jimmy X. Huang
arXiv:2111. 15255v2 Announce Type: replace-cross Abstract: The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations.
By Zongmin Liu
arXiv:2610.01191v1 Announce Type: new
Abstract: An optical character recognition(OCR) system can scan paper and extract text, making people's jobs easier. While numerous OCR systems are accessible in...
By Faias Satter, Noor Masrur, Sk. Md. Masudul Ahsan
arXiv:2208. 00335v5 Announce Type: replace Abstract: Rule extraction is a central problem in interpretable machine learning because it seeks to convert opaque predictive behavior into human-readable symbolic structure.
By Caleb Princewill Nwokocha
arXiv:2607. 10548v1 Announce Type: cross Abstract: Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clustering.
By Zhihao Yao, Yuxuan Gu, Jixuan Yin, Bo Li
arXiv:2509. 21160v2 Announce Type: replace-cross Abstract: With the growing use of large language models, concerns over content authenticity have spurred a variety of watermarking schemes.
By Soham Bonnerjee, Subhrajyoty Roy, Sayar Karmakar
Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction using machine learning. One of these subfields is...
arXiv:2609.07606v1 Announce Type: new
Abstract: Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction...
By Johann Faouzi
The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.
By Danel Slabbert, Simon Malan, Herman Kamper
The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.
By Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri
Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.
Automating the classification and extraction of PII from emails using AWS The post Build and Run an Intelligent Document Processing (IDP) System in the Cloud appeared first on Towards Data Science .
By Thomas Reid