arXiv:2605. 06605v2 Announce Type: replace Abstract: Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.
By Shai Feldman, Yaniv Romano
arXiv:2609.38914v1 Announce Type: new
Abstract: Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare a...
By Priyanath Maji, Spandan Ghose Chowdhury
arXiv:2609.13149v1 Announce Type: new
Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all...
By Aditya Karnam Gururaj Rao, Arjun Jaggi
arXiv:2609.35760v2 Announce Type: replace-cross
Abstract: When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The age...
By Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
arXiv:2607. 17545v1 Announce Type: new Abstract: Language agents depend on memory across interactions.
By Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan
arXiv:2609.15309v1 Announce Type: new
Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...
By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung
The paper introduces RareTrap, a framework that estimates the probability of severe behaviors in black‑box large language models. RareTrap constructs a geometry‑aware mapping from a low‑dimensional latent space into token‑embedding space using a surrogate LLM, creating an explicit and reproducible distribution over input prompts. By applying a response‑level performance function and sequential rare‑event simulation, RareTrap concentrates evaluations on increasingly severe behaviors while preserving probability, enabling estimation of such behaviors with as few as 200 evaluations across multiple open‑weight and frontier models.
By Elsayed Eshra, Ali Al-Lawati, Dongwon Lee, Suhang Wang
arXiv:2601. 21522v2 Announce Type: replace-cross Abstract: The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials.
By Sagi Meir, Tommer D. Keidar, Noam Levi, Shlomi Reuveni, Barak Hirshberg
The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.
By Liner Xiang, Wenbo Zhang, Hengrui Cai
arXiv:2606. 08696v1 Announce Type: cross Abstract: Counterfactual recourse aims to provide actionable feature changes that would alter an unfavorable decision made by a predictive model.
By Yasuo Tabei
arXiv:2605.25893v2 Announce Type: replace
Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monit...
By Aoxi Liu, Yupeng Chen, James Oldfield, Guanzhe Hong, Junchi Yu, Baoyuan Wu, Philip Torr, Adel Bibi
arXiv:2607. 10694v1 Announce Type: cross Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device.
By Thomas Tsouparopoulos, Iordanis Koutsopoulos