The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.
By Changyi Li, Yu Xiao
arXiv:2607. 02307v1 Announce Type: cross Abstract: Several SLOG test categories explicitly involve directional distinctions (modifier position shifts, argument extraction positions), yet AM-Parser, the previous SOTA, uses an AM algebra whose operations do not encode direction.
By Zichao Wei
arXiv:2609.21509v1 Announce Type: new
Abstract: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How m...
By Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella, Samy Bengio
The paper evaluates two large language models, Claude Sonnet 4.5 and Claude Opus 5, on the bidirectional English Resource Grammar (ERG) tasks of generating English from Minimal Recursion Semantics (MRS) and parsing English into MRS. In generation, Opus achieves 76.3 BLEU—surpassing a 72k‑pair trained system and matching a million‑pair system—while Sonnet scores 65.7 BLEU, rising to 69.6 when selecting from ACE’s candidates. In parsing, both models lag behind ACE, attaining only 57.2 and 65.5 F₁ respectively, with exact‑match on about 1 % of sentences, highlighting that high generation scores do not guarantee accurate semantic parsing.
By Soham Dan
arXiv:2607. 18961v1 Announce Type: new Abstract: Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix.
By Remo Pareschi
The paper investigates why Transformers struggle more with structural than lexical compositional generalisation. It argues that this disparity stems from low structural type diversity rather than an inherent limitation of Transformers. By creating linguistically diverse variants of the COGS and SLOG datasets, the authors show that type diversity correlates equally with generalisation in both lexical and structural cases, challenging previous explanations of the difficulty.
By Anssi Moisio, Mathias Creutz, Mikko Kurimo
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
By Haris Riaz, Hyungji Kim, Mihai Surdeanu
arXiv:2606. 25450v1 Announce Type: new Abstract: Traditional evaluations measure a learning algorithm's final performance on an i.
By Jinghan Zhang, Zerui Cheng, Shiqi Chen, Ge Zhang, Wenhao Huang, Jiashuo Liu, Junxian He, Tianle Cai
The paper proposes a new way to evaluate compositional generalization by examining which structural or lexical identifications allow held‑out COGS examples to be considered admissible based on training data. Sentences are modeled as functors from syntactic addresses to lexical tokens, and selective collapses induce Kan extensions that propagate observed associations. Across 21 COGS generalization types, admissibility follows distinct identification profiles, while residual failures highlight unsupported structural templates, providing data‑side diagnoses of what the training corpus licenses without training a predictive model.
By Akihiro Maeda, Thomas Seiller, Yohei Oseki
arXiv:2608. 10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step.
By I\c{s}{\i}l \"Ozg\"u, Yaoxuan Wu, Guy Van den Broeck, Miryung Kim
arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.
By Alexandre Bouayad
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
By Joshua S. Schiffman