Hugging Face Trending Papers

Exposure is Optional: Learning Unlike Coordination in Language Models

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.

arXiv AI
Sep 18

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances explores whether a language model can embed a signal in natural language that another independent instance can detect without shared memory or coordination training. The study introduces a cooperative signalling game where a Sender describes two words, one hidden, and a Receiver must identify the target. Seven contemporary models from four architectural families were tested on 300 word pairs, revealing that most struggle to coordinate when signals must be undetectable, though one frontier model performs near-perfectly even after filtering, and that models can also use this capability for deliberate misdirection.

By Alexander Shirnin, Aleksey Kudelya
arXiv Computation and Language
Sep 18

Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.

By Giuseppe Samo, Vivi Nastase, Paola Merlo
arXiv Computation and Language
Aug 28

Compositional Generalization via Structural Identification in a Category-Theoretic Framework

The paper proposes a new way to evaluate compositional generalization by examining which structural or lexical identifications allow held‑out COGS examples to be considered admissible based on training data. Sentences are modeled as functors from syntactic addresses to lexical tokens, and selective collapses induce Kan extensions that propagate observed associations. Across 21 COGS generalization types, admissibility follows distinct identification profiles, while residual failures highlight unsupported structural templates, providing data‑side diagnoses of what the training corpus licenses without training a predictive model.

By Akihiro Maeda, Thomas Seiller, Yohei Oseki
arXiv Computation and Language
Aug 31

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.

By Titus von der Malsburg, Sebastian Pad\'o
arXiv AI
3d ago

NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space

The paper introduces NinaXander, a method for composing frozen language models from different architecture families by inserting a trained shared‑latent adapter between their layers. By running the initial layers of one model, converting the intermediate representation with the adapter, and then continuing with the remaining layers of another model, multiple composed models can be created without retraining. Experiments with RWKV and Pythia show that while some compositions preserve syntactic quality and reduce memory usage, none match the parent model’s accuracy and language‑modeling performance drops on out‑of‑domain data.

By Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri
arXiv AI
Sep 7

Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation

The paper studies structural priming in language model production by conducting controlled sentence‑completion experiments on dative constructions. Results show that language models exhibit priming effects, especially when sentences are semantically coherent, with stronger relative increases for double‑object datives and larger absolute increases for prepositional‑object datives. The study also finds that primed completions involve more lexico‑semantic repetition, indicating that priming operates across syntactic, lexical, and semantic levels.

By Giulia Pucci, Ruizhe Li, Arabella Sinclair
arXiv Computation and Language
Sep 17

Modelling Adjectival Modification Effects on Semantic Plausibility

The paper investigates how adjectival modifiers affect the semantic plausibility of events, using the Adept benchmark of 16,000 English sentence pairs that differ by a single adjective. Experiments show that sentence transformers, despite being conceptually suited to the task, underperform compared to models like RoBERTa. The authors provide an error analysis and discuss the implications of their findings for future work on balancing training and test data.

By Anna Golub, Beate Zywietz, Annerose Eichel
arXiv Computation and Language
Sep 3

Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.

By Yixuan Wang, Freda Shi, Kanishka Misra