WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.
By Tyler McDonald, Ali Emami
arXiv:2607. 20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formulations as interchangeable and conflates reasoning errors with interface failures.
By Sagnik Nath, Edith Aurora Graf, Liang Zhang, Diego Zapata-Rivera
arXiv:2608. 12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings.
By Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
arXiv:2602.00377v3 Announce Type: replace
Abstract: Existing knowledge probing methods rely on pre-defined queries, limiting extraction to known concepts. We introduce DecompressionLM, a stateless fr...
By Zhaochen Hong, Jiaxuan You
SyntaxBench is a diagnostic benchmark and statistical evaluation framework for character‑level reasoning in large language models, comprising five core tasks—character counting, letter containment, palindrome detection, edit distance, and longest‑string selection—and a harder substring‑extraction stress test called index_to_span. The benchmark uses paired English and random‑string inputs, zero‑, one‑, and four‑shot prompts, and evaluates models from 2B to 32B parameters across multiple reasoning modes. It reports a wide range of metrics, including exact‑match and relaxed accuracy, Cohen’s kappa, McNemar tests, bootstrap confidence intervals, Kendall’s tau, class‑conditional metrics, tokenization analysis, and multiple‑comparison‑corrected tests.
By Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)
arXiv:2609.25001v1 Announce Type: new
Abstract: Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, a...
By Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
arXiv:2607. 22652v1 Announce Type: new Abstract: Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA).
By Yike Wu, Nan Hu, Guilin Qi, Guohui Xiao, Chen Jiang, Xinchun Zou, Yuchen Lu, Songlin Zhai, Yongrui Chen, Yuyang Zhang, Xiaoguang Li, Lifeng Shang, Jiaoyan Chen, Jeff Z. Pan
arXiv:2508. 10971v2 Announce Type: replace-cross Abstract: Knowledge graphs (KGs) can be enhanced through rule mining; however, the resulting logical rules are often difficult for humans to interpret due to their inherent complexity and the idiosyncratic labeling conventions of individual KGs.
By Nasim Shirvani-Mahdavi, Chengkai Li
arXiv:2607. 08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them.
By Siddhartha Jain, Ameya Velingker
arXiv:2608.30437v1 Announce Type: new
Abstract: Graph-augmented large language models often assume that graph evidence produced by external computation and placed in the input can be used by the nati...
By Xiaoyu Guo, Pengcheng Chen, Jiong Yu, Yi Lu, Yaohua Wang, Ziyang Li
The paper introduces the Abstraction Agent, a zero‑shot pipeline that employs a large language model to automatically generate continuous strategic features from a natural‑language game description, score private states, and cluster them into abstraction buckets without any game‑specific evaluators or training data. The pipeline consists of four phases—feature discovery with calibration anchors, batched private‑state scoring, correlation‑based feature selection, and k‑means clustering—and achieves significant reductions in lifted‑strategy exploitability in heads‑up no‑limit Texas hold’em and outperforms scalar rank baselines in ROVER Trials. The method also transfers to other games such as four‑card Pot‑Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, demonstrating that it can uncover strategic concepts that align with recognized game theory insights.
By Boning Li, Longbo Huang
arXiv:2607. 00527v1 Announce Type: new Abstract: Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime.
By Zhiyue Xu, Fandi Meng, Kaijie Xu, Clark Verbrugge, Simon Lucas, Jian Zhao