The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.
By Javier Aguilar Mart\'in
The paper studies how a code‑world model can be perfectly accurate on the portion of the state space that a sampling gate can observe while potentially being arbitrarily wrong elsewhere. By treating the unobservable interior as an annular freeze mode, the authors formalize the notion of reach and show that acceptance with certainty fixes the model only on the reachable query set, leaving the rest as a gauge. Experiments on a minimal ring instrument demonstrate that a single channel width parameter can move the model through regimes of being unfalsifiable and harmless, falsifiable and costly, or instantly falsified, illustrating how topology relative to reach governs danger, repair, and mitigation strategies.
By Javier Aguilar Mart\'in
arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
arXiv:2607. 14169v1 Announce Type: new Abstract: Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches over.
By Javier Aguilar Mart\'in
The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv:2607. 23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B".
By Paul Simpson, John Kozak, Lisa Doake