arXiv AI By Guangzhe Zhang

Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

Read the original on arXiv AI →

The study examines how context projection—replacing older tool observations with concise, addressable excerpts—affects performance in ReVerPi, a Pi extension that archives observations and matches full to projected continuations. Across 86 source‑reading runs and 641 model requests, 15 completed pairs achieved identical success rates (12/15 per arm), while 12 boundary runs halted when the first arm failed, revealing that projection can reduce logical tokens by 25% but increase median pair tokens by 29% and total suffix requests from 35 to 55. The analysis highlights the importance of retaining all intervention boundaries, executing both arms independently, and reporting completion, token usage, and interaction details to accurately assess stopping rules and resource aggregation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv AI
Sep 7

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

The paper investigates whether the object selected in a grounded language‑model pipeline actually reaches the reader, a failure that can break the handoff between stages. By auditing 600 HybridQA questions across three selector families, the authors find that exact key lookup and title matching recover the selected object in all 1,463 resolvable records, but body‑only BM25 omits it in 26.6% of cases at cutoff five, while hybrid retrieval with reranking omits it only 1.0%. The study also shows that misalignment between selected and retrieved objects can reduce exact match scores by up to 31 points, and introduces the Returned‑Object Profile (ROP) as a tool for reproducible auditing.

By Siddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu