arXiv Computation and Language

An Incomplete Loop: Deductive, Inductive, and Abductive Reasoning in Language Models

arXiv AI
Aug 28

Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning

The paper investigates whether large language models (LLMs) follow Occam's Razor when performing inductive and abductive reasoning. It introduces a synthetic framework for generating questions that require both types of reasoning and a new automated metric to evaluate the simplicity and correctness of generated hypotheses. Experiments show that while LLMs can handle simple scenarios, they struggle with complex world models and producing high‑quality, simplest hypotheses, even when using advanced reasoning techniques.

By Yunxin Sun, Abulhair Saparov
arXiv AI
Sep 2

Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs

The paper investigates how large language models balance instruction-following with pattern completion when the two objectives conflict. By creating dialogues where a user instruction to act in a target way T is opposed by assistant turns that demonstrate a competing pattern P, the authors measure instruction-following rates across 13 models and 16 instructions over up to 50 turns. Results show wide variability (1%–99%) in instruction adherence, with robustness influenced by instruction content, output format, and chain-of-thought reasoning, but overall instruction-following remains brittle under induction pressure.

By Carolina Camassa, Derek Shiller
arXiv AI
Sep 3

Thinking effort aligns between humans and reasoning models in abductive reasoning

The study examines how the effort expended by large reasoning models (LRMs) compares to that of humans during abductive reasoning tasks. By analyzing reaction times and reasoning traces, the authors find that LRMs and humans exhibit similar patterns of effort and error types. They also demonstrate that decoding strategies allowing models to explore multiple reasoning paths further align the models’ reasoning costs with human effort.

By Henry Arthur
arXiv Computation and Language
Sep 4

LLMs Learn Better In-Context from Rules than from Examples

The paper investigates how large language models learn new tasks in-context, comparing rule-based instruction following to example-based few-shot prompting across five diverse tasks. Results show that models generally learn more reliably from rule descriptions than from examples alone, and adding more examples does not consistently improve performance. Instruction tuning further enhances rule-based learning while preserving example-based capabilities, with rule advantages being strongest for algebraic tasks and weaker for tasks requiring distributional sensitivity or parametric knowledge.

By Xiang Fu, Seungmin Cho, Yukyung Lee, Najoung Kim
arXiv AI
Aug 18

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

arXiv:2608. 14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery.

By Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban
arXiv Machine Learning
Jun 5

Reasoning Models Don't Just Think Longer, They Move Differently

arXiv:2605. 15454v2 Announce Type: replace-cross Abstract: Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory.

By Anders Gj{\o}lbye, Lars Kai Hansen, Sanmi Koyejo
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen