The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.
By Javier Aguilar Mart\'in
The paper investigates the safety of Code World Models, where a language model generates an executable world model that a planner uses, and the model is accepted if it reproduces sampled transitions. It defines the pipeline’s danger as the expected risk, showing that the probability of missing a critical event across N independent rollouts is (1‑r)^N, and that an additional acceptance sample adds to the exponent. Experiments on hybrid instruments reveal that mode‑blind models can be exploited, and the authors provide theoretical bounds on localization budgets and demonstrate that acceptance only guarantees sample consistency, covering about two percent of the planner’s queries.
arXiv:2608. 10986v1 Announce Type: cross Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops.
By Nicol\'as Vera Z\'u\~niga
The paper investigates a new failure mode of tool‑augmented large language model agents: calling non‑existent tools with arguments that do not match any declared schema. It introduces a five‑class taxonomy of tool hallucination, presents a training‑free closed‑world resolver that checks registry membership and signatures, and demonstrates that hallucinations persist across ten hosted models and various invocation surfaces, including the Model Context Protocol. The authors release a Hallucinated‑Tools Benchmark to enable comparison of resolver methods.
By Laxmipriya Ganesh Iyer
arXiv:2607. 07665v1 Announce Type: new Abstract: Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules.
By Shiheng Zhang
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan