arXiv AI By Juliusz Ziomek, William Bankes, Lorenz Wolf, Shyam Sundhar Ramesh, Xiaohang Tang, Ilija Bogunovic

LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?

Read the original on arXiv AI →

arXiv:2602. 16902v4 Announce Type: replace Abstract: We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 14

LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?

LLM‑BabyBench transforms the BabyAI gridworld into a fully observable, purely textual setting that isolates planning as the sole source of failure. By serialising the entire grid, providing formal instructions, and validating actions deterministically, the benchmark introduces the PPD suite—Predict, Plan, and Decompose tasks—each scored with metrics that separate mission understanding from sequencing. Across a range of large language models, simulation accuracy is high while planning success drops sharply beyond a model‑specific horizon, revealing that plan length—not grid size—drives failure and that models often commit to a single corridor‑shaped route without backtracking.

By Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou
arXiv AI
Sep 17

WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.

By Tyler McDonald, Ali Emami
arXiv AI
Aug 19

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

G-ReAct is a reasoning framework that frames deep search as state evolution over a fixed-topology query graph, enabling explicit tracking of search progress and constraint preservation. It generates high-quality trajectories for fine-tuning and provides structured guidance during inference without extra fine-tuning. Experiments show that with only 1.9K generated trajectories, a Qwen3 model achieves strong accuracy on BrowseComp-ZH and XBench, outperforming larger open-source baselines, and consistently improves existing LLMs on deep-search tasks.

By Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
arXiv AI
Sep 24

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model. "whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."

By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini