arXiv AI

InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information

InteractBench is a new benchmark that evaluates large language models on interactive competitive programming problems, where key information is revealed only through queries to an interactor. It contains 322 curated problems from Codeforces, AtCoder, IOI, and ICPC, each with executable local interactors for offline testing. The study finds that even top reasoning models struggle with these tasks, highlighting gaps in information acquisition, state tracking, and protocol adherence.

arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
arXiv AI
Aug 13

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

arXiv:2608. 12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation.

By Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
arXiv Machine Learning
1d ago

Online Verification of Language Model Responses Under Cost Constraints

The paper introduces OMVV, an online multi-verifier algorithm that maintains a pool of weak verifiers with varying costs and performance. It adaptively selects a verifier each round using an online score combiner and exponential-weights routing, providing distribution-free guarantees on false-accept and false-reject rates. Experiments on reasoning benchmarks show OMVV achieves higher accuracy at lower verification cost than any single fixed verifier across different budgets.

By Erfan Hajihashemi, Yanning Shen
arXiv AI
Sep 10

ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language

ProcArena is a new benchmark for evaluating large language models on natural‑language to PL/SQL translation tasks. It contains 3,998 executable tasks across 157 databases, covering nine development subscenarios in PostgreSQL and Oracle, and supports both direct generation and interactive multi‑turn scenarios. Experiments on seven models show that even the best performers achieve only about 62% accuracy in direct mode and 58% in interactive mode, highlighting the difficulty of realistic NL‑to‑PL/SQL development.

By Hang Zhang, Chaokun Wang, Yuzhi Pan, Ziyao Zhong, Shuo Cao, Yue Xue, Zeyu Huang, Xingwei Zhou, Fang Niu, Bofan Xie, Guanchen Ge, Leqi Zheng, Ziyang Liu, Xiannian Cao, Pengcheng Ge