Sequential Bayesian Evaluation of Large Language Model Behavior
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, s...
arXiv:2609.24890v1 Announce Type: cross Abstract: Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with f...
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome.
The paper introduces RareTrap, a framework that estimates the probability of severe behaviors in black‑box large language models. RareTrap constructs a geometry‑aware mapping from a low‑dimensional latent space into token‑embedding space using a surrogate LLM, creating an explicit and reproducible distribution over input prompts. By applying a response‑level performance function and sequential rare‑event simulation, RareTrap concentrates evaluations on increasingly severe behaviors while preserving probability, enabling estimation of such behaviors with as few as 200 evaluations across multiple open‑weight and frontier models.
arXiv:2606. 01182v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning.