arXiv AI

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?

arXiv:2602. 16902v4 Announce Type: replace Abstract: We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs).

arXiv Computation and Language
Sep 14

LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?

LLM‑BabyBench transforms the BabyAI gridworld into a fully observable, purely textual setting that isolates planning as the sole source of failure. By serialising the entire grid, providing formal instructions, and validating actions deterministically, the benchmark introduces the PPD suite—Predict, Plan, and Decompose tasks—each scored with metrics that separate mission understanding from sequencing. Across a range of large language models, simulation accuracy is high while planning success drops sharply beyond a model‑specific horizon, revealing that plan length—not grid size—drives failure and that models often commit to a single corridor‑shaped route without backtracking.

By Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou
arXiv AI
Sep 17

WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.

By Tyler McDonald, Ali Emami
arXiv AI
Aug 19

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

G-ReAct is a reasoning framework that frames deep search as state evolution over a fixed-topology query graph, enabling explicit tracking of search progress and constraint preservation. It generates high-quality trajectories for fine-tuning and provides structured guidance during inference without extra fine-tuning. Experiments show that with only 1.9K generated trajectories, a Qwen3 model achieves strong accuracy on BrowseComp-ZH and XBench, outperforming larger open-source baselines, and consistently improves existing LLMs on deep-search tasks.

By Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
arXiv AI
Sep 24

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model. "whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."

By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
arXiv AI
Aug 26

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

arXiv:2603.16654v3 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espe...

By Xiaojie Gu, Sherry T. Tong, Aosong Feng, Sophia Simeng Han, Jinghui Lu, Yingjian Chen, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Rex Ying, Irene Li
arXiv Machine Learning
Jun 5

IR3DE: A Linear Router for Large Language Models

arXiv:2606. 06098v1 Announce Type: cross Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialized tasks via domain-expert LLMs.

By Eros Fan\`i, O\u{g}uzhan Ersoy
arXiv AI
Sep 18

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.

By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic
arXiv AI
Jun 15

Fractured Chain-of-Thought Reasoning

arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.

By Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong