arXiv Computation and Language

Compositionality and the lexicon in evolutionary semantics

Hugging Face Trending Papers
Jun 25

Compositionality and the lexicon in evolutionary semantics

Formal semantics has shown that sentence meanings arise by recursively composing lexical meanings, yet much of the literature on semantic universals models either lexicons with fixed signal structures or holistic composition without interpretable lexical parts. We introduce a framework that integrates this fundamental insight of formal semantics in evolutionary modeling, by allowing lexical meanings and a composition function to co-evolve under pressures for conceptual simplicity and communicative accuracy.

arXiv Machine Learning
Jun 18

From Mechanistic to Compositional Interpretability

arXiv:2605. 08934v2 Announce Type: replace Abstract: Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components.

By Ward Gauderis, Thomas Dooms, Steven T. Homer, Kola Ayonrinde, Geraint A. Wiggins
arXiv Computation and Language
4d ago

Type Diversity Enables Transformers to Generalise Compositionally

The paper investigates why Transformers struggle more with structural than lexical compositional generalisation. It argues that this disparity stems from low structural type diversity rather than an inherent limitation of Transformers. By creating linguistically diverse variants of the COGS and SLOG datasets, the authors show that type diversity correlates equally with generalisation in both lexical and structural cases, challenging previous explanations of the difficulty.

By Anssi Moisio, Mathias Creutz, Mikko Kurimo
arXiv Computation and Language
2d ago

Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics

The paper investigates Mandarin Chinese reduplicative constructions that repeat two-character base words or their constituents, such as expressions meaning ‘in good health’ or ‘discuss a bit’. Using Tencent word embeddings, the study demonstrates that distributional semantics can recover known semantic and grammatical properties of these reduplications, revealing clear semantic and pragmatic differentiation between the two patterns. Procrustes analysis shows that the overall organization of the base-word space is largely preserved in the reduplication space, with local mismatches indicating discourse-pragmatic reorganization.

By Chaoyi Wu, Yu-Hsiang Tseng, R. Harald Baayen
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.

arXiv AI
Jun 3

SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory

arXiv:2511. 16275v4 Announce Type: replace-cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language models (LLMs) in safety-critical scenarios, as it enables them to abstain from responding when uncertain, thereby avoiding hallucinations, i.

By Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu