arXiv AI

Generating Concept Lexicalizations via Dictionary-Based Cross-Lingual Sense Projection

arXiv:2604. 14397v2 Announce Type: replace-cross Abstract: We study the task of automatically expanding WordNet-style lexical resources to new languages through sense generation.

arXiv AI
Sep 2

Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages

Inspicio is an open‑vocabulary pipeline that links tokens in historical or low‑resource languages to synsets in the Open English WordNet without needing a source‑language sense inventory. It uses an instruction‑tuned LLM to generate two English translations, candidate dictionary definitions, and English lemmas, then performs hybrid retrieval combining dense definition similarity, sparse lemma matching, and Maximal Marginal Relevance re‑ranking. Evaluated on Latin, Ancient Greek, PREMOVE, and Italian data, the best configuration achieves 96% Recall@50 on a perception‑verb test set and remains competitive in out‑of‑domain and cross‑lingual scenarios.

By Michele Ciletti
arXiv Computation and Language
Sep 28

Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries

The paper tackles the problem of automatically generating dictionary definitions for learner’s dictionaries, focusing on simplicity and clarity. It introduces a new evaluation framework that uses large language models as judges, validated against human annotators with comparable agreement levels. The authors also present an iterative simplification approach that produces definitions scoring highly on their criteria and exhibiting lexical simplicity.

By Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay, Taro Watanabe
arXiv Computation and Language
Sep 24

MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

MetaHOPE is an error‑severity‑aware annotation framework designed to evaluate how well machine translation (MT) and large language models (LLMs) translate metaphors. The authors applied MetaHOPE to three state‑of‑the‑art systems—GoogleMT, GPT5.4, and Hunyuan‑7b—using two human‑annotated metaphor corpora (VUAMC and PSUCMC) for English‑to‑Chinese and Chinese‑to‑English translation. They also produced a bilingual post‑edited gold reference, creating a new resource for metaphor translation research.

By Jiahui Liang, Lifeng Han
arXiv Computation and Language
Aug 28

Cross-lingual Representation Learning via Centroid Intervention Fusion

The paper introduces Centroid Intervention Fusion (CIF), a framework that merges multiple multilingual intervention projections into a single language-shared operator for inference-time modification of large language models. CIF improves cross-lingual transfer without updating model parameters and achieves up to +3.378 percentage points better performance than prior pairwise intervention baselines across several benchmarks, including low-resource languages. The authors provide code at https://github.com/VRCMF/CIF.git.

By Wei Sun, Marie-Francine Moens
arXiv Computation and Language
1d ago

Latent Core Tokenizer: Compress, but Meaningfully

The paper introduces the Latent Core Tokenizer (LCT), a language‑agnostic method that first discovers reusable linguistic units using Minimum Description Length, entropy‑based boundary signals, and morphotactic constraints before building a shared vocabulary. With a 200K‑token vocabulary across 104 languages, LCT shows lower fertility and higher MorphScore than BPE, Unigram, and parity‑aware BPE, while keeping tokenization cost comparable across languages. On four multilingual downstream benchmarks, LCT outperforms the baselines by 1.48, 1.83, and 2.00 aggregate points, demonstrating that compression alone does not guarantee representation quality and underscoring the role of morphology‑driven structural discovery.

By Felermino D. M. A. Ali, Millicent Ochieng, Ogbemi Ekwejunor-Etchie, Ade Famoti, Jacki O'Neill, Debjit Paul