arXiv AI

ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization

arXiv:2607. 05185v1 Announce Type: new Abstract: Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence.

arXiv Machine Learning
Sep 22

Subgoal Search For Complex Reasoning Tasks

arXiv:2108.11204v4 Announce Type: replace-cross Abstract: Humans excel in solving complex reasoning tasks through a mental process of moving from one idea to a related one. Inspired by this, we propo...

By Konrad Czechowski, Tomasz Odrzyg\'o\'zd\'z, Marek Zbysi\'nski, Micha{\l} Zawalski, Krzysztof Olejnik, Yuhuai Wu, {\L}ukasz Kuci\'nski, Piotr Mi{\l}o\'s
arXiv Machine Learning
Sep 11

A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.

By Deblina Kar
arXiv AI
Sep 17

Clueing up LLMs with Tool-Augmented Deductive Reasoning

The paper introduces a text-based, multi-agent version of the board game Clue to test multi-step deductive reasoning in large language models (LLMs). Six LLM-based agents (GPT‑4o‑mini and Gemini‑2.5‑Flash) play turn‑based games, and a tool‑augmented approach uses a structured possibility matrix to convert implicit game state into explicit remaining possibilities, thereby offloading memory and deductive constraints from the agents. The study compares this tool‑augmented method against a baseline to assess its impact on reasoning quality and task success in a strategic reasoning environment.

By Rebecca Ansell, Autumn Toney-Wails
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao
arXiv AI
Sep 24

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.

By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv AI
Sep 17

What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

The paper presents a systematic mapping of recent chess research involving humans, engines, neural and reinforcement‑learning systems, large language models (LLMs), and hybrid approaches. It identifies 84 core study families and classifies them by agent type, strategic‑reasoning stages, and evaluation dimensions, highlighting a strong focus on situation assessment, evaluation, and action selection while noting gaps in planning, explanation, metacognition, and human–AI collaboration. The study also distinguishes hybrid systems by integration timing and cautions that improved human performance in evaluations does not automatically imply human–AI synergy.

By Paolo Ciancarini, Remo Pareschi