arXiv AI
Sep 17

WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.

By Tyler McDonald, Ali Emami
arXiv AI
Jul 24

Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

arXiv:2607. 20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures.

By Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera
arXiv AI
Aug 14

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

arXiv:2608. 12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings.

By Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
arXiv AI
3d ago

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

SyntaxBench is a diagnostic benchmark and statistical evaluation framework for character‑level reasoning in large language models, comprising five core tasks—character counting, letter containment, palindrome detection, edit distance, and longest‑string selection—and a harder substring‑extraction stress test called index_to_span. The benchmark uses paired English and random‑string inputs, zero‑, one‑, and four‑shot prompts, and evaluates models from 2B to 32B parameters across multiple reasoning modes. It reports a wide range of metrics, including exact‑match and relaxed accuracy, Cohen’s kappa, McNemar tests, bootstrap confidence intervals, Kendall’s tau, class‑conditional metrics, tokenization analysis, and multiple‑comparison‑corrected tests.

By Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)
arXiv Computer Vision
Sep 22

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

arXiv:2609.25001v1 Announce Type: new Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...

By Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan