Hugging Face Trending Papers

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models.

arXiv AI
Sep 15

Thought without systematicity? Evaluating reasoning models on rule induction tasks

The paper investigates whether current reasoning models exhibit systematicity—the idea that understanding one concept should extend to closely related variations—by extending rule induction tasks from cognitive science. Using task isomorphisms like recombination and substitution, the authors generate structurally equivalent task variants and test models on them. Results show that while models can solve the original tasks, they frequently fail on these equivalent variants, indicating a lack of systematicity in their reasoning abilities.

By Simon Schug, Brenden M. Lake
arXiv Machine Learning
Sep 10

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

The paper investigates whether reasoning representations—explanations for large language model outputs—aid humans in evaluating those outputs. A controlled human study tested six reasoning formats across tasks of varying complexity, measuring structural understanding, error detection, and trust calibration. Results revealed a mismatch: participants favored planning- and decomposition-based representations, yet simpler chain-of-thought traces better supported verification, trust, and interpretability, while preferred formats increased calibration risks.

By Jaewoo Lim, Sungbok Shin, Sanghyun Hong
arXiv AI
Sep 3

Thinking effort aligns between humans and reasoning models in abductive reasoning

The study examines how the effort expended by large reasoning models (LRMs) compares to that of humans during abductive reasoning tasks. By analyzing reaction times and reasoning traces, the authors find that LRMs and humans exhibit similar patterns of effort and error types. They also demonstrate that decoding strategies allowing models to explore multiple reasoning paths further align the models’ reasoning costs with human effort.

By Henry Arthur