arXiv AI By Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Read the original on arXiv AI →

arXiv:2608. 13547v1 Announce Type: new Abstract: LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

The study examines how context projection—replacing older tool observations with concise, addressable excerpts—affects performance in ReVerPi, a Pi extension that archives observations and matches full to projected continuations. Across 86 source‑reading runs and 641 model requests, 15 completed pairs achieved identical success rates (12/15 per arm), while 12 boundary runs halted when the first arm failed, revealing that projection can reduce logical tokens by 25% but increase median pair tokens by 29% and total suffix requests from 35 to 55. The analysis highlights the importance of retaining all intervention boundaries, executing both arms independently, and reporting completion, token usage, and interaction details to accurately assess stopping rules and resource aggregation.

By Guangzhe Zhang
arXiv AI
Sep 15

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

IBBench-Light is a paired evaluation framework that tests language models on both executing procedures and reading text from the same external record. The benchmark uses twelve semantic bases to generate 144 matched pairs per model, with four instruction‑quantized models producing 1,152 greedy responses. Metrics such as Paired Exact‑Contract Accuracy (PECA) reveal that models like Qwen achieve high success on individual prompts but only 97 complete pairs, highlighting the importance of paired evaluation.

By Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao
arXiv AI
Jul 24

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya