arXiv Machine Learning

What Can a Recurrent State Safely Forget?

arXiv:2609. 23366v1 Announce Type: new Abstract: Recurrent models must preserve information that changes future behavior while suppressing hidden-state error.

arXiv AI
Aug 18

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

The paper investigates whether a small, directly addressable change in the hidden state of a learned world model can steer its future predictions along a desired counterfactual trajectory. Using a 192‑dimensional recurrent model in a two‑object collision setting, the authors identify low‑rank latent carriers—specifically a rank‑4 patch—that, when applied, successfully redirect a 12‑step autonomous rollout without further intervention. The study demonstrates that this compact intervention interface consistently works across independently trained checkpoints and intervention times, while various control experiments confirm the specificity of the effect.

By Yang Liu, Yuming Chen
arXiv Computer Vision
Sep 25

BARRIER: Bounded Activation Regions for Robust Information Erasure

BARRIER (Bounded Activation Regions for Robust Information Erasure) is a method for machine unlearning that confines parameter updates to a controlled activation space, allowing stronger erasure of targeted concepts while limiting collateral damage to other representations. By employing interval arithmetic, it derives a closed‑form bound on worst‑case representation changes in protected regions, which serves as a knowledge‑preservation objective. The approach is architecture‑agnostic, compatible with existing erasure objectives, and empirically shows competitive performance in both classification and generative tasks, with enhanced robustness against adversarial recovery attacks.

By Jan Miksa, Patryk Krukowski, Przemys{\l}aw Spurek, Dawid Damian Rymarczyk, Marcin Sendera
arXiv AI
Sep 10

How to Verify Probabilistic Consistency of Predictive Models

The paper presents an interactive probabilistically checkable proof (PCP) protocol that allows a polynomial‑time verifier to check the approximate consistency of a probabilistic predictor defined by two circuits, P and Q. By evaluating these circuits at a few points and querying a proof oracle that encodes a witnessing probability distribution, the verifier can confirm that the predictor’s many conditional‑probability claims are self‑consistent. The authors also establish that the problem of verifying l₂‑approximate consistency for explicit probabilistic claims lies in NP, with certificates of size O(mn + log B), and show how to eliminate dependence on the input bit‑precision B through a small additive gap.

By Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser
Hugging Face Trending Papers
Aug 18

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates an executable world model that a planner uses, and the model is accepted if it reproduces sampled transitions. It defines the pipeline’s danger as the expected risk, showing that the probability of missing a critical event across N independent rollouts is (1‑r)^N, and that an additional acceptance sample adds to the exponent. Experiments on hybrid instruments reveal that mode‑blind models can be exploited, and the authors provide theoretical bounds on localization budgets and demonstrate that acceptance only guarantees sample consistency, covering about two percent of the planner’s queries.