Learning Patterns and Abstractions from Perceptual Sequences
arXiv:2503. 10973v2 Announce Type: replace Abstract: Cognition swiftly breaks high-dimensional sensory streams into familiar parts and uncovers their relations.
The paper reviews Cobweb, a computational model of categorization and concept formation, and extends it to include chunks and their acquisition. It introduces rellis/, an implementation that applies this unified theory to learning context-free grammars, demonstrating the system’s ability to represent syntactic knowledge, parse and generate sentences, and learn compositional structures from sample parses. The authors discuss related work on concepts and chunks and suggest directions for future research.
arXiv:2503. 10973v2 Announce Type: replace Abstract: Cognition swiftly breaks high-dimensional sensory streams into familiar parts and uncovers their relations.
The paper proposes that large language models (LLMs) encode high‑level concepts as linear directions within their activation space and that they can use subspaces and vector algebra to perform tasks. By analyzing functional modules and residual streams during in‑context learning (ICL), the authors find that LLMs can create evidence‑accumulating subspaces and solve ICL tasks through simple algebraic operations within those subspaces.
arXiv:2607. 18961v1 Announce Type: new Abstract: Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix.
The paper surveys computational models that use programs as representations for concepts, arguing that programs could serve as a universal language for concepts. It highlights how humans learn and generalize from sparse data by expressing knowledge in rich structural formats, and evaluates how program-based models contribute to this goal. The authors propose that adopting programs as a universal representational language could enhance concept learning across diverse domains.
arXiv:2603. 01227v3 Announce Type: replace Abstract: We propose the Lattice Representation Hypothesis of large language models: a symbolic backbone that grounds conceptual hierarchies and logical operations in embedding geometry.
arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.
arXiv:2602. 02886v3 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) promote interpretability by grounding predictions in human-understandable concepts.
arXiv:2606.05087v2 Announce Type: replace Abstract: Frequent verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'make a...
arXiv:2604. 26157v4 Announce Type: replace-cross Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations.
The paper studies how transformer representations evolve across layers by examining the intrinsic dimensionality (ID) of token embeddings and their neighborhood structures. It finds that closed‑class tokens expand and collapse earlier than open‑class tokens, and that these changes are linked to shifts in local geometry. The authors compare encoder and decoder models, showing distinct layer‑wise behaviors, and demonstrate that geometric features alone can predict a token’s part‑of‑speech and reveal how semantic content changes across layers.
The paper investigates how verb–noun decomposition, a common strategy for recognizing assembly actions, generalizes to novel combinations of familiar components. Through a systematic study on three datasets (MECCANO, HAViD, and IMPACT), the authors find that while decomposition avoids the zero‑probability ceiling of atomic classifiers, its performance still heavily depends on the co‑occurrence patterns seen during training. The analysis reveals that errors concentrate on the larger‑vocabulary component, that shared‑encoder training can entangle components and worsen generalization, and that these issues stem from primitive support, vocabulary asymmetry, and component entanglement.
The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.