Hugging Face Trending Papers

Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet

This paper proposed an algorithm for part-of-speech (POS) tagging senses of a bilingual dictionary. The algorithm is applied on the Al-Mawrid Arabic-English dictionary.

arXiv Computation and Language
Sep 28

Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries

The paper tackles the problem of automatically generating dictionary definitions for learner’s dictionaries, focusing on simplicity and clarity. It introduces a new evaluation framework that uses large language models as judges, validated against human annotators with comparable agreement levels. The authors also present an iterative simplification approach that produces definitions scoring highly on their criteria and exhibiting lexical simplicity.

By Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay, Taro Watanabe
arXiv Computation and Language
Sep 3

MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

MUDIDI is a two-stage framework designed to digitize multilingual dictionaries that are currently only available as scanned images. The first stage assesses character recognition and markup preservation, while the second stage segments dictionary entries and maps them into the SIL Multi-Dictionary Formatter schema. The authors also release a dataset of 30 annotated dictionaries and benchmark OCR, LLM, and VLM systems, finding that LLMs generally outperform others and that providing additional context improves digitization quality.

By David Setiawan, Temuulen Khishigsuren, Milind Agarwal, Pagnarith Pit, Aso Mahmudi, Ekaterina Vylomova
arXiv AI
Sep 2

Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages

Inspicio is an open‑vocabulary pipeline that links tokens in historical or low‑resource languages to synsets in the Open English WordNet without needing a source‑language sense inventory. It uses an instruction‑tuned LLM to generate two English translations, candidate dictionary definitions, and English lemmas, then performs hybrid retrieval combining dense definition similarity, sparse lemma matching, and Maximal Marginal Relevance re‑ranking. Evaluated on Latin, Ancient Greek, PREMOVE, and Italian data, the best configuration achieves 96% Recall@50 on a perception‑verb test set and remains competitive in out‑of‑domain and cross‑lingual scenarios.

By Michele Ciletti
arXiv Computation and Language
Sep 4

Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models

The paper evaluates large language models (LLMs) on Arabic morphosyntactic tagging and dependency parsing, a challenging task due to rich morphology and orthographic ambiguity. It compares zero‑shot prompting with retrieval‑based in‑context learning across pre‑tokenized, raw‑text, and cascaded settings, finding that relevant demonstrations significantly boost performance. The best LLMs nearly match supervised systems but need extensive annotated data for demonstrations and high computational resources. All code and data are publicly released.

By Mohamed Adel, Bashar Alhafni, Nizar Habash
arXiv AI
Aug 25

The Multilingual FrameNet Corpus

The paper presents the Multilingual FrameNet Corpus (mFNC), a resource that expands the English Berkeley FrameNet by integrating and harmonizing language‑specific corpora in nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian, and Swedish. Experiments with various model architectures on mFNC consistently surpass existing state‑of‑the‑art Frame Semantic Parsers in both multilingual and cross‑lingual scenarios, highlighting the value of multilingual training data. The mFNC and the trained Frame Semantic Parser models are publicly released on GitHub.

By Beatrice Fiuman\`o, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi