arXiv:2610.03137v1 Announce Type: new
Abstract: Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in...
By Florian Strohm, Patrick Wagner, Jannik Schwab, Marco Huber
arXiv:2609.37378v1 Announce Type: cross
Abstract: Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occur...
By Hossein Resani, Javen Qinfeng Shi
JevAdvBench introduces the first adversarial benchmark for reinforcement‑learning‑based calibrated decision (RLCD) models, providing 812 typed questions across 66 scenarios and a black‑box attack suite of 9,744 single‑edit variants. The benchmark evaluates attacks by comparing each perturbed decision to the model’s own clean decision and to an identical re‑run, revealing that rewording changes decisions by only 1.2 percentage points while certain injected opinions can flip 12.1% of decisions and lower confidence below 0.8 in 38% of cases. These findings demonstrate that RLCD models can be significantly misled by seemingly innocuous input edits, underscoring the need to treat the state as untrusted in applications.
By Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang
arXiv:2609.25686v1 Announce Type: cross
Abstract: Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent...
By Chenyu Zhang, Wonbin Kweon, Jiawei Han
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan
arXiv:2603.12717v2 Announce Type: replace-cross
Abstract: Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are de...
By Tuan Duong Trinh, Basim Azam, Mohammed Ishaq Ansari, Mohammed Yaqoob Ansari, Naveed Akhtar
arXiv:2609.08123v1 Announce Type: cross
Abstract: A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly t...
By Suyog Khanal, Arun Kumar A V, Santu Rana
arXiv:2604. 11840v3 Announce Type: replace-cross Abstract: Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises.
By Sandro Andric
The study examines how renaming option labels in typed decision models affects model behavior. By swapping the names of two options (e.g., from 0/1 to no/yes) while keeping the underlying rubrics unchanged, the authors observed a dramatic shift in decision rankings—AUC dropped from .94 to .23 and answer flips increased by 70.4 per hundred. The effect is amplified with more options and depends on the semantic polarity of the labels, yet the models still maintain a zero type‑error rate.
By Yu Sun, Junhao Xu, Jiajia Shi, Zijin Yang
arXiv:2607. 07196v1 Announce Type: cross Abstract: Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world, and returning a success or safety verdict.
By Christian Oefinger, Finn Rasmus Sch\"afer, Korbinian Moller, Mattia Piccinini, Johannes Betz
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen