Hugging Face Trending Papers

Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords

The paper introduces the Fixed Suffix Dependency Ratio (FSDR) as a metric to measure how much loanwords depend on fixed derivational suffixes for gender assignment. Analyzing 1,832 Latvian noun lemmas, it finds that feminine loanwords rely more on fixed suffixes, whereas masculine loanwords are more often assigned freely, a pattern that has intensified in recent usage. This asymmetry highlights a dual-track mechanism in gender assignment under language contact.

arXiv Computation and Language
Sep 4

Fixed Suffix Dependency Ratio: Quantifying the Dual-Track Mechanism of Gender Assignment in Latvian Loanwords

The paper introduces the Fixed Suffix Dependency Ratio (FSDR) as a metric to measure how much loanwords depend on fixed derivational suffixes for gender assignment. Analyzing 1,832 Latvian noun lemmas, it finds that feminine loanwords rely more on fixed suffixes while masculine loanwords are largely free‑choice, a pattern that has intensified in recent usage. FSDR offers a quantitative tool for distinguishing morphological anchoring from default gender in language contact scenarios.

By Yelingyun Zhang, Atis Kapenieks, Marina Platonova
arXiv Computation and Language
4d ago

Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations

The paper investigates gender bias in machine translation evaluation metrics using an occupation-balanced subset of GAMBIT+ across seven English‑source language pairs, including a new German extension. It finds that masculine translations tend to receive higher scores and that biases align with stereotypical gender representations, though the strength varies by evaluator and language. The study highlights that assessing bias requires multiple dimensions beyond a single aggregate measure.

By Orfeas Menis Mastromichalakis, Giorgos Filandrianos, Wafaa Mohammed, Giuseppe Attanasio, Chrysoula Zerva
arXiv Computation and Language
Sep 17

The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models

The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.

By Elzbieta Dawidek (University of Lower Silesia DSW Ideis)
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv AI
Sep 11

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

The study investigates how speech‑to‑speech (S2S) models handle gender, distinguishing between the acoustic voice and the content’s gender cues. Experiments across five models in English, Spanish, and Mandarin show that while the rendered voice remains unbiased, the models consistently attribute speaker gender based on textual content rather than voice. When content and voice disagree, misgendering rates soar to 90%, whereas agreement yields only 2% misgendering.

By Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji