arXiv AI

Fuzzy Segmentations of a String

The article addresses a specific data clustering problem: identifying groups of adjacent text segments of a suitable length that match a fuzzy pattern defined by a sequence of fuzzy properties. It proposes a heuristic algorithm that uses a prefix structure to efficiently map text segments to fuzzy properties, and proves that for the special case of unit-length segments (fuzzy string matching), the algorithm finds all matching segments. Additionally, it presents a dynamic programming approach to determine the best segmentation of an entire text based on a fuzzy pattern.

arXiv Computation and Language
Sep 15

Recovering the Zipfian Distribution in Unsupervised Term Discovery

The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.

By Danel Slabbert, Simon Malan, Herman Kamper
arXiv Machine Learning
Aug 28

Absolute indices for determining compactness, separability and number of clusters

The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.

By Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri