Holographic Neural PCFG for Unsupervised Parsing
arXiv:2607. 08063v1 Announce Type: cross Abstract: Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone.
Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone. Recent neural parameterizations of PCFGs achieve strong performance in both supervised and unsupervised parsing, yet rely on high-capacity black-box networks for rule scoring -- as exemplified by the Neural PCFG family -- leaving rule probabilities without an interpretable mathematical form.
arXiv:2607. 08063v1 Announce Type: cross Abstract: Unsupervised constituency parsing aims to accurately induce latent tree structures from raw text alone.
The paper investigates how different representations of Korean constituency structure affect parsing performance. It compares three formats—Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS—derived from the Penn Korean Treebank, using gold segmentation and labels to evaluate transition-based parsers. Results show that fine-grained morphological and XPOS information yields the best parsing accuracy, while eojeol-based representations offer shorter transition sequences but lower performance when only UPOS is used.
CWoMP (Contrastive Word‑Morpheme Pretraining) is a new approach for generating interlinear glossed text that treats morphemes as atomic form‑meaning units with learned representations. It uses a contrastively trained encoder to align words in context with their constituent morphemes in a shared embedding space, and an autoregressive decoder that retrieves morpheme sequences from a mutable lexicon of these embeddings. The method yields interpretable predictions grounded in lexicon entries and allows users to improve results at inference time by expanding the lexicon without retraining, achieving superior performance and efficiency on diverse low‑resource languages, especially in extremely low‑resource settings.
arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.
arXiv:2604. 26157v4 Announce Type: replace-cross Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations.
arXiv:2608.23448v1 Announce Type: new Abstract: This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradi...
The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, using a discriminative conditional random field over a word lattice built from dictionaries. It addresses data scarcity by generating over two million sentences with large language models. Experiments show the method surpasses traditional morphological analyzers and neural sequence models, achieving 99.62% target word reading accuracy and very low phoneme error rates on the Joyo‑Kanji‑Yomi benchmark.
arXiv:2510. 19698v3 Announce Type: replace Abstract: Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional rule learning.
The paper introduces a context‑aware neural grapheme‑to‑phoneme (G2P) system for unsegmented languages like Japanese, combining a discriminative conditional random field over a dictionary‑based word lattice with large language model‑generated training data. By generating over two million synthetic sentences, the method addresses data scarcity and achieves superior performance compared to traditional morphological analyzers and neural sequence models. On the Joyo‑Kanji‑Yomi benchmark, it attains 99.62% target word reading accuracy, 0.32% target word phoneme error rate, and 0.14% sentence phoneme error rate.
ThaiTrees is a 342‑million‑token corpus of Thai text spanning news, Wikipedia, spoken transcripts, and social media, automatically parsed under the Universal Dependencies framework. The authors provide a reproducible pipeline for cleaning, processing, and parsing the data, and release the resulting CoNLL‑U files and a frequency lexicon in machine‑readable formats. This resource enables researchers to search grammatical relations and study syntactic distributions at scale.
arXiv:2610.01519v1 Announce Type: cross Abstract: Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified...
The paper tackles syntactic parsing for Urdu, a morphologically rich language, and reports state‑of‑the‑art results for both constituency and dependency parsing. It introduces four key contributions: converting the CLE‑UTB phrase structure treebank into a dependency treebank with language‑specific mapping rules, a novel sequence labeling scheme that unifies the parsing task, training contextualized word representations on a 220‑million‑token Urdu corpus, and a parsing framework that employs both single‑task and multi‑task learning. Experiments show that the multi‑task setup boosts performance, achieving an F1 score of 91.39 for constituency parsing and a labeled attachment score of 85.69 for dependency parsing, improving over previous results.