arXiv:2608.04186v3 Announce Type: replace
Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (L...
By Mullosharaf K. Arabov, Saidali M. Pirzoda, Behruz A. Sultonov
arXiv:2604. 14397v2 Announce Type: replace-cross Abstract: We study the task of automatically expanding WordNet-style lexical resources to new languages through sense generation.
By David Basil, Chirooth Girigowda, Bradley Hauer, Sahir Momin, Ning Shi, Grzegorz Kondrak
This paper proposed an algorithm for part-of-speech (POS) tagging senses of a bilingual dictionary. The algorithm is applied on the Al-Mawrid Arabic-English dictionary.
The paper investigates why Word-in-Context (WiC) remains difficult for language models, suggesting that the lack of an explicit sense inventory contributes to the challenge. By evaluating open LLMs on both WiC and traditional Word Sense Disambiguation (WSD) tasks, the authors find that providing candidate senses—akin to WSD—consistently improves WiC performance. Human evaluation indicates that many WiC errors stem from label ambiguity or mismatched sense boundaries, with models often over‑discriminating senses and making overly fine‑grained distinctions.
By Yi Zhou, Kiamehr Rezaee, Danushka Bollegala, Mohammad Taher Pilehvar, Jose Camacho-Collados
MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.
By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova
KoNeoBench is a curated dataset designed to evaluate large language models’ understanding of Korean neologisms. It contains 1,785 recently attested Korean words from online news since 2020, each accompanied by usage examples, word‑formation analyses, and dictionary‑style definitions. The authors define four evaluation tasks, report results from recent models and a human baseline, and find that current LLMs struggle with recovering source components, distinguishing semantic categories, and generating accurate definitions.
By Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam
Word Sense Disambiguation has advanced rapidly for English and a handful of well-resourced modern languages, but it continues to assume the existence of a sense inventory and a word-to-sense mapping i...
Inspicio is an open‑vocabulary pipeline that links tokens in historical or low‑resource languages to synsets in the Open English WordNet without needing a source‑language sense inventory. It uses an instruction‑tuned LLM to generate two English translations, candidate dictionary definitions, and English lemmas, then performs hybrid retrieval combining dense definition similarity, sparse lemma matching, and Maximal Marginal Relevance re‑ranking. Evaluated on Latin, Ancient Greek, PREMOVE, and Italian data, the best configuration achieves 96% Recall@50 on a perception‑verb test set and remains competitive in out‑of‑domain and cross‑lingual scenarios.
By Michele Ciletti
The paper introduces a reverse sign‑language dictionary that recognizes signs from continuous signing without relying on gloss labels. It does this by first captioning a sign‑level video clip into a free‑form procedural description using an open‑weight vision‑language model, then retrieving the closest entry from a multilingual sentence encoder’s vocabulary of target descriptions. Experiments on a Japanese Sign Language dialogue corpus show that fine‑tuning the captioner boosts seen‑class retrieval from 4.5 % to 49 % and improves unseen‑class retrieval from 11.5 % to 21 %, approaching the performance of a standard closed‑set classifier while enabling open‑vocabulary recognition.
By Santiago Poveda-Guti\'errez, Hideki Nakayama, Mayumi Bono
CWoMP (Contrastive Word‑Morpheme Pretraining) is a new approach for generating interlinear glossed text that treats morphemes as atomic form‑meaning units with learned representations. It uses a contrastively trained encoder to align words in context with their constituent morphemes in a shared embedding space, and an autoregressive decoder that retrieves morpheme sequences from a mutable lexicon of these embeddings. The method yields interpretable predictions grounded in lexicon entries and allows users to improve results at inference time by expanding the lexicon without retraining, achieving superior performance and efficiency on diverse low‑resource languages, especially in extremely low‑resource settings.
By Morris Alper, Enora Rice, Bhargav Shandilya, Alexis Palmer, Lori Levin
arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.
By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
SignTrace is a system that tackles the reverse‑lookup problem in Chinese sign language by allowing users to describe a sign in natural language and retrieve the corresponding entry from a dictionary of 6,699 signs. It combines large‑language‑model–based dictionary enrichment, action extraction, seven‑channel retrieval, and candidate reranking, achieving a 94.0% Hit@1 and a mean reciprocal rank of 0.9540 on a benchmark of 500 queries. The system has been deployed for user trials, receives positive informal feedback, and processes median queries in 13.37 seconds.
By Zengji Tu, Xingye Zhu, Ningjing Wang, Tingyi Huang, Yangjunfeng Zhu, Dai Wan