arXiv Computation and Language

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

The paper critiques the normalized edit distance metric used for evaluating lexicons derived from unsupervised word discovery, noting its bias toward large clusters and its failure to account for the distribution of true classes across clusters. It proposes two new metrics—one that weights cluster size when measuring within‑cluster consistency and another that evaluates how true words are spread across clusters—drawing on clustering theory. Experiments on synthetic and real‑world lexicons show that these combined metrics better correlate with ground‑truth distributions and are more robust to evaluation biases.

arXiv Computation and Language
Sep 15

Recovering the Zipfian Distribution in Unsupervised Term Discovery

The paper investigates unsupervised term discovery in speech, comparing centre-based clustering methods like K‑means with graph‑based clustering using the Leiden algorithm. It finds that graph clustering produces lexicons whose type frequencies follow a Zipfian distribution, outperforming K‑means, GMM, and BIRCH across word‑ and syllable‑level discovery in three languages. Agglomerative clustering with average linkage also performs well but is less efficient and offers less control over the distribution.

By Danel Slabbert, Simon Malan, Herman Kamper
arXiv Computation and Language
Aug 21

SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit

arXiv:2608. 19472v1 Announce Type: new Abstract: Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing.

By Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale, Dirk Speelman
arXiv Machine Learning
Sep 17

TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation

TACTICS is a method for selecting evaluation samples in machine translation that explicitly optimizes for coverage of rare linguistic categories, document-level coherence, and distributional fidelity to the full corpus. It builds a hierarchical taxonomy from a locale style guide, classifies segments, and chooses a fixed-budget subset that better represents the full range of phenomena a system must handle. Compared to random, lexical, or embedding-based selection, TACTICS improves coverage of rare categories and yields more accurate system rankings with fewer segments.

By Prasanth Bathala, Anubhav Shrimal, Sukhdeep Singh Kharbhanda, Pradyumna Lanka, Rohit Dhaipule
arXiv Computation and Language
5d ago

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux