arXiv AI

Emergent Latent-State Computation under Stochastic Volatility

arXiv:2607. 25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks.

arXiv Machine Learning
4d ago

HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling

HALO introduces a hyperspherical VAE to constrain continuous latent representations to a fixed‑radius shell, stabilizing numerical fluctuations. It then employs a masked autoregressive model that balances parallel decoding with temporal correlation learning, reducing inference steps and improving stability. Experiments show HALO achieves state‑of‑the‑art generation performance with significantly better inference efficiency compared to existing baselines.

By Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang
arXiv AI
Jun 19

How Transparent is DiffusionGemma?

arXiv:2606. 20560v1 Announce Type: cross Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors.

By Joshua Engels, Callum McDougall, Bilal Chughtai, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue, Jo\~ao Gabriel Lopes de Oliveira, Rohin Shah, Neel Nanda
arXiv Machine Learning
5d ago

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.

By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv Machine Learning
Jul 2

TRIE: An Evaluation Framework for Stochastic PDE Surrogates

arXiv:2607. 00196v1 Announce Type: new Abstract: Many scientific systems exhibit uncertainty from stochastic forcing, unresolved degrees of freedom, or imperfect observations, making reliable surrogate forecasting fundamentally distributional rather than pointwise.

By Bharat Srikishan, Javier E. Santos, Nikhil Muralidhar, Charles D. Young
arXiv Machine Learning
Jun 29

Textual Belief States for World Models: Identifiable Representation Learning Under Strict Mediation

arXiv:2606. 27681v1 Announce Type: new Abstract: World models in partially observed environments rely on latent representations that summarize interaction history, but in many modern LLM-based architectures predictive performance fails to reflect representation quality due to history bypass, rendering the latent state unidentifiable.

By Xiang Gao, Kaiwen Dong, Yuguang Yao, Padmaja Jonnalagedda, Kamalika Das