Hugging Face Trending Papers

Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet

Read the original on Hugging Face Trending Papers →

This paper proposed an algorithm for part-of-speech (POS) tagging senses of a bilingual dictionary. The algorithm is applied on the Al-Mawrid Arabic-English dictionary.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 28

Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries

The paper tackles the problem of automatically generating dictionary definitions for learner’s dictionaries, focusing on simplicity and clarity. It introduces a new evaluation framework that uses large language models as judges, validated against human annotators with comparable agreement levels. The authors also present an iterative simplification approach that produces definitions scoring highly on their criteria and exhibiting lexical simplicity.

By Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay, Taro Watanabe
arXiv Computation and Language
Sep 3

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.

By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova