arXiv Machine Learning

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

arXiv Machine Learning
Aug 26

Replicable Conformal Prediction

arXiv:2608.23638v1 Announce Type: cross Abstract: Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration...

By Marios Papamichalis, Regina Ruane, Theofanis Papamichalis
arXiv AI
Aug 28

Invocation-Level Reliability of Tool-Using Agents

The paper investigates the reliability of tool‑using agents, focusing on two failure modes: selecting the wrong tool and constructing incorrect arguments. It introduces a correct‑invocation rate metric to distinguish these errors and evaluates five open‑weight models on multi‑step tasks up to depth 8, finding that by depth 6 about 70% of a model’s clean‑context capability is lost due to earlier mistakes. The study reveals that exact‑match scoring against a fixed gold trajectory forces severity and recovery parameters to extreme values, and proposes a conditional‑on‑state scoring remedy that yields more realistic severity estimates.

By Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee, Abhijit Dasgupta
arXiv AI
Aug 19

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.

By Javier Aguilar Mart\'in