The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.
By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta
arXiv:2606. 19636v1 Announce Type: cross Abstract: Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal.
By Luca Zhou, Sajel Shah, Emanuele Rodol\`a, Roberto Dess\`i
The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.
By Guiv Farmanfarmaian
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
arXiv:2601. 21522v2 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials.
By Sagi Meir, Tommer D. Keidar, Noam Levi, Shlomi Reuveni, Barak Hirshberg
arXiv:2607. 28576v1 Announce Type: cross Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought.
By Iliya Mirzaei