arXiv AI

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.

Hugging Face Trending Papers
Aug 18

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates an executable world model that a planner uses, and the model is accepted if it reproduces sampled transitions. It defines the pipeline’s danger as the expected risk, showing that the probability of missing a critical event across N independent rollouts is (1‑r)^N, and that an additional acceptance sample adds to the exponent. Experiments on hybrid instruments reveal that mode‑blind models can be exploited, and the authors provide theoretical bounds on localization budgets and demonstrate that acceptance only guarantees sample consistency, covering about two percent of the planner’s queries.

arXiv Machine Learning
Aug 31

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

The paper studies how a code‑world model can be perfectly accurate on the portion of the state space that a sampling gate can observe while potentially being arbitrarily wrong elsewhere. By treating the unobservable interior as an annular freeze mode, the authors formalize the notion of reach and show that acceptance with certainty fixes the model only on the reachable query set, leaving the rest as a gauge. Experiments on a minimal ring instrument demonstrate that a single channel width parameter can move the model through regimes of being unfalsifiable and harmless, falsifiable and costly, or instantly falsified, illustrating how topology relative to reach governs danger, repair, and mitigation strategies.

By Javier Aguilar Mart\'in
arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
Aug 26

When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs

The paper introduces Evidence-Carrying Termination (ECT), a method that allows tool‑using large language model agents to declare a task complete only when a typed certificate links every required answer claim to valid, in‑scope trace evidence and a deterministic replay confirms the claimed value. In controlled experiments across 48 synthetic tasks and 576 trajectories, ECT eliminated unsafe completions and premature unsupported terminations, outperforming existing termination critics and controllers while maintaining comparable supported completion rates.

By Jason Liu
arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu