Identifiability Without Gaussianity: Symbolic World Models and Near-Infinite Temporal Consistency
arXiv:2606. 12471v2 Announce Type: replace-cross Abstract: Klindt, LeCun, and Balestriero (arXiv:2605.
arXiv:2606. 10934v1 Announce Type: new Abstract: A common assumption holds that enough observational and interventional data, given to a strong enough predictor, suffices.
arXiv:2606. 12471v2 Announce Type: replace-cross Abstract: Klindt, LeCun, and Balestriero (arXiv:2605.
arXiv:2603.05335v3 Announce Type: replace-cross Abstract: Modern predictive systems combine predictors, sequential monitors, prediction sets, and online strategies, each with a different certificate...
arXiv:2607. 15629v1 Announce Type: cross Abstract: Topos causal models recast causal inference inside a topos: a causal world is a presheaf, an intervention is a characteristic map into the subobject classifier, and reasoning is carried out in the intuitionistic internal language.
The paper introduces a framework that distinguishes world models by the channel they represent—environment, agent, or joint agent‑environment—using computational mechanics to define canonical predictive models as ε-transducers or ε-machines. It shows how closed‑loop coupling induces support‑restricted models whose states factor through the joint causal states, and demonstrates with a POMDP example that such restriction can reduce an otherwise infinite‑state environment model to a finite one.
The paper investigates why differentiable causal discovery methods that encode expert priors as forbidden-edge constraints via an Augmented Lagrangian (ALM) penalty—termed the "guide, not bind" approach—often fail. It identifies two key failures: (1) the sequential penalty‑ramping ALM suppresses a true edge before counterfactual checks can detect it, and the proposed adaptive relaxation rule DADU violates necessary conditions for safe relaxation, leading to a high failure rate across thousands of training runs; (2) the standard correlation‑matching objective inherently ties a true edge and its reverse to the same cost, whereas covariance matching can separate them by a provable margin. The authors provide theoretical propositions, corollaries, and empirical evidence to support these claims.
arXiv:2608. 15645v1 Announce Type: new Abstract: Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference.
arXiv:2608. 07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong.
arXiv:2606. 04421v1 Announce Type: new Abstract: Many current agentic systems and LLM pipelines correct mistakes by optimizing outcome reward.
arXiv:2608. 14004v1 Announce Type: new Abstract: In-context learning is commonly formalized as inference from examples of a function.
arXiv:2609.25388v1 Announce Type: cross Abstract: A classical question in statistics is which observable quantities to condition on when drawing inferences about unobservable targets. For conformal p...
arXiv:2608. 15147v1 Announce Type: new Abstract: Machine intelligence has conquered the symbolic world but stalled at the physical one.
The paper identifies a single direction in the unembedding matrix of large language models that encodes the unigram distribution of the training corpus, acting as a Bayesian prior when the model is uncertain. By projecting the final prediction state onto this direction, the authors derive a per‑token prior loading factor, λ, which decreases as context becomes more informative and decomposes predictions into a tempered prior and a context‑driven likelihood. Experiments across four model families (Llama, Qwen, Gemma, Pythia) show that larger models rely less on the prior in high‑context settings and that manipulating λ can steer predictions toward or away from the unigram prior in KL divergence.