arXiv AI By Yu He, Yingxi Li, Colin White, Ellen Vitercik

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

Read the original on arXiv AI →

arXiv:2505. 24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
arXiv Machine Learning
Jun 25

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

arXiv:2606. 24965v1 Announce Type: cross Abstract: Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned knowledge to problem instances that are harder than those seen in training.

By Anirban Das, Joanne Boisson, Irtaza Khalid, Sumita Garai, Steven Schockaert
arXiv AI
Sep 18

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

PetriBench is a compact, fully self‑contained, and scalable benchmark that evaluates large language model (LLM) reasoning over dynamic state spaces using Petri nets. It organizes reasoning into four task families with Easy, Medium, and Hard levels, each generated by increasing structural complexity and evaluated against exact ground truth. Experiments across proprietary and open‑weight models show that accuracy consistently drops with difficulty, revealing distinct task‑specific capability profiles, while test‑time compute and procedural generation affect performance differently across tasks.

By Pyrros Koussios, Benjamin J\"ager, John Hua Yao, Ajay Sridhar, Violet Xiang, Chenhao Li