arXiv AI By Simon Schug, Brenden M. Lake

Thought without systematicity? Evaluating reasoning models on rule induction tasks

Read the original on arXiv AI →

The paper investigates whether current reasoning models exhibit systematicity—the idea that understanding one concept should extend to closely related variations—by extending rule induction tasks from cognitive science. Using task isomorphisms like recombination and substitution, the authors generate structurally equivalent task variants and test models on them. Results show that while models can solve the original tasks, they frequently fail on these equivalent variants, indicating a lack of systematicity in their reasoning abilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis

The paper critiques current systematic generalization benchmarks for oversimplifying the task by relying on linear action composition, productivity-based tests, and action-explicit goals. It introduces TranSGrid, a new testbed that integrates deductive, inductive, and abductive reasoning in a unified setting. Experiments with seven Transformers show a significant performance drop on TranSGrid compared to standard held-out tests, indicating that existing simplifications mask the true difficulty of systematic generalization.

By Chengwen Qi, Deheng Ye, Yatao Bian
arXiv Machine Learning
Sep 11

A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.

By Deblina Kar
Hugging Face Trending Papers
Jun 11

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models.