arXiv Computation and Language By Simon Malan, Danel Slabbert, Herman Kamper

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

Read the original on arXiv Computation and Language →

The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 15

Recovering the Zipfian Distribution in Unsupervised Term Discovery

The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.

By Danel Slabbert, Simon Malan, Herman Kamper