arXiv AI

Structural Generalization on SLOG without Hand-Written Rules

arXiv:2604. 26157v4 Announce Type: replace-cross Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations.

arXiv Computer Vision
2d ago

Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition

The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.

By Changyi Li, Yu Xiao
arXiv Computation and Language
Sep 25

Scoring Both Directions: LLMs realize the MRS they cannot reliably parse

The paper evaluates two large language models, Claude Sonnet 4.5 and Claude Opus 5, on the bidirectional English Resource Grammar (ERG) tasks of generating English from Minimal Recursion Semantics (MRS) and parsing English into MRS. In generation, Opus achieves 76.3 BLEU—surpassing a 72k‑pair trained system and matching a million‑pair system—while Sonnet scores 65.7 BLEU, rising to 69.6 when selecting from ACE’s candidates. In parsing, both models lag behind ACE, attaining only 57.2 and 65.5 F₁ respectively, with exact‑match on about 1 % of sentences, highlighting that high generation scores do not guarantee accurate semantic parsing.

By Soham Dan
arXiv Computation and Language
Sep 14

Type Diversity Enables Transformers to Generalise Compositionally

The paper investigates why Transformers struggle more with structural than lexical compositional generalisation. It argues that this disparity stems from low structural type diversity rather than an inherent limitation of Transformers. By creating linguistically diverse variants of the COGS and SLOG datasets, the authors show that type diversity correlates equally with generalisation in both lexical and structural cases, challenging previous explanations of the difficulty.

By Anssi Moisio, Mathias Creutz, Mikko Kurimo
arXiv Computation and Language
Aug 28

Compositional Generalization via Structural Identification in a Category-Theoretic Framework

The paper proposes a new way to evaluate compositional generalization by examining which structural or lexical identifications allow held‑out COGS examples to be considered admissible based on training data. Sentences are modeled as functors from syntactic addresses to lexical tokens, and selective collapses induce Kan extensions that propagate observed associations. Across 21 COGS generalization types, admissibility follows distinct identification profiles, while residual failures highlight unsupported structural templates, providing data‑side diagnoses of what the training corpus licenses without training a predictive model.

By Akihiro Maeda, Thomas Seiller, Yohei Oseki
arXiv Machine Learning
Jun 25

Weave of Formal Thought

arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.

By Alexandre Bouayad