Hugging Face Trending Papers

Holographic Neural PCFG for Unsupervised Parsing

Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone. Recent neural parameterizations of PCFGs achieve strong performance in both supervised and unsupervised parsing, yet rely on high-capacity black-box networks for rule scoring -- as exemplified by the Neural PCFG family -- leaving rule probabilities without an interpretable mathematical form.

arXiv Computation and Language
Aug 28

Representing and Parsing Korean Constituency Structure at Different Levels of Granularity

The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.

By Jungyeul Park, KyungTae Lim, Zihao Huang, Eunkyul Leah Jo, Yige Chen, Chulwoo Park
arXiv Computation and Language
Sep 2

CWoMP: Morpheme Representation Learning for Interlinear Glossing

CWoMP (Contrastive Word‑Morpheme Pretraining) is a new approach for generating interlinear glossed text that treats morphemes as atomic form‑meaning units with learned representations. It uses a contrastively trained encoder to align words in context with their constituent morphemes in a shared embedding space, and an autoregressive decoder that retrieves morpheme sequences from a mutable lexicon of these embeddings. The method yields interpretable predictions grounded in lexicon entries and allows users to improve results at inference time by expanding the lexicon without retraining, achieving superior performance and efficiency on diverse low‑resource languages, especially in extremely low‑resource settings.

By Morris Alper, Enora Rice, Bhargav Shandilya, Alexis Palmer, Lori Levin
arXiv Machine Learning
Jun 25

Weave of Formal Thought

arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.

By Alexandre Bouayad
arXiv Computation and Language
Sep 18

Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, using a discriminative conditional random field over a word lattice built from dictionaries. It addresses data scarcity by generating over two million sentences with large language models. Experiments show the method surpasses traditional morphological analyzers and neural sequence models, achieving 99.62% target word reading accuracy and very low phoneme error rates on the Joyo‑Kanji‑Yomi benchmark.

By Rui Hu, Zhenpeng Zhan, Xiaolong Lin
Hugging Face Trending Papers
Sep 17

Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, combining a discriminative conditional random field over a dictionary‑based word lattice with large language model‑generated training data. By generating over two million synthetic sentences, the method addresses data scarcity and achieves superior performance compared to traditional morphological analyzers and neural sequence models. On the Joyo‑Kanji‑Yomi benchmark, it attains 99.62% target word reading accuracy, 0.32% target word phoneme error rate, and 0.14% sentence phoneme error rate.

arXiv Computation and Language
Sep 24

ThaiTrees: Thai Syntactic Dependency Trees Across Domains

ThaiTrees is a 342‑million‑token corpus of Thai text spanning news, Wikipedia, spoken transcripts, and social media, automatically parsed under the Universal Dependencies framework. The authors provide a reproducible pipeline for cleaning, processing, and parsing the data, and release the resulting CoNLL‑U files and a frequency lexicon in machine‑readable formats. This resource enables researchers to search grammatical relations and study syntactic distributions at scale.

By Attapol T. Rutherford, Papatchol Thientong
arXiv AI
2d ago

Auto-Formalizing Neuro-Symbolic Predictors

arXiv:2610.01519v1 Announce Type: cross Abstract: Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified...

By Samuele Bortolotti, Weixin Chen, Han Zhao, Andrea Passerini, Stefano Teso, Antonio Vergari
arXiv AI
Sep 25

Multi-Task Learning by using Contextualized Word Representations for Syntactic Parsing of a Morphologically Rich Language

The paper tackles syntactic parsing for Urdu, a morphologically rich language, and reports state‑of‑the‑art results for both constituency and dependency parsing. It introduces four key contributions: converting the CLE‑UTB phrase structure treebank into a dependency treebank with language‑specific mapping rules, a novel sequence labeling scheme that unifies the parsing task, training contextualized word representations on a 220‑million‑token Urdu corpus, and a parsing framework that employs both single‑task and multi‑task learning. Experiments show that the multi‑task setup boosts performance, achieving an F1 score of 91.39 for constituency parsing and a labeled attachment score of 85.69 for dependency parsing, improving over previous results.

By Toqeer Ehsan, Miriam Butt, Sarmad Hussain, Hassan Alhuzali, Ali Al-Laith