arXiv AI By Mahnoor Shahid, Hannes Rothe

ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization

Read the original on arXiv AI →

arXiv:2607. 05185v1 Announce Type: new Abstract: Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Subgoal Search For Complex Reasoning Tasks

arXiv:2108.11204v4 Announce Type: replace-cross Abstract: Humans excel in solving complex reasoning tasks through a mental process of moving from one idea to a related one. Inspired by this, we propo...

By Konrad Czechowski, Tomasz Odrzyg\'o\'zd\'z, Marek Zbysi\'nski, Micha{\l} Zawalski, Krzysztof Olejnik, Yuhuai Wu, {\L}ukasz Kuci\'nski, Piotr Mi{\l}o\'s
arXiv Machine Learning
Sep 11

A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

The paper introduces a multi-stage rule‑chaining framework for the Abstraction and Reasoning Corpus (ARC), aiming to model cognitive generalization by inferring abstract rules from few examples. It combines three solvers—a deterministic rule discovery module, a pattern‑composition engine, and a structural abstraction layer—executed sequentially in a fallback hierarchy that reuses earlier reasoning traces to improve interpretability and generalization. The system achieved over 95% accuracy on ARC tasks, demonstrating strong performance across deterministic, compositional, and abstract categories.

By Deblina Kar
arXiv AI
Sep 17

Clueing up LLMs with Tool-Augmented Deductive Reasoning

The paper introduces a text-based, multi-agent version of the board game Clue to test multi-step deductive reasoning in large language models (LLMs). Six LLM-based agents (GPT‑4o‑mini and Gemini‑2.5‑Flash) play turn‑based games, and a tool‑augmented approach uses a structured possibility matrix to convert implicit game state into explicit remaining possibilities, thereby offloading memory and deductive constraints from the agents. The study compares this tool‑augmented method against a baseline to assess its impact on reasoning quality and task success in a strategic reasoning environment.

By Rebecca Ansell, Autumn Toney-Wails
arXiv AI
Sep 18

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers is a cooperative framework that jointly adapts planning, execution, and evaluation for tool-integrated reinforcement learning agents. It consists of a Planning Player that generates tasks, an Execution Player that creates multi-turn trajectories with Python tool calls, and an Evaluation Player that builds executable verifiers, all coordinated by role‑specific rewards under GRPO. The approach outperforms prior baselines on mathematical and general reasoning benchmarks and yields a verifier with high adversarial detection accuracy and more discriminative reward signals.

By Wenjie Liao, Liangjie Zhao, Zehong Cao