xWhyL: Causal Interactive Learning
arXiv:2609.26037v1 Announce Type: new Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for lear...
arXiv:2508. 11214v2 Announce Type: replace-cross Abstract: Explanations of cognitive behavior often appeal to computations over representations.
arXiv:2609.26037v1 Announce Type: new Abstract: Explanations are central to causal reasoning, and cognitive science has long established that the human drive to explain is itself a mechanism for lear...
arXiv:2411. 08875v4 Announce Type: replace Abstract: Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to find them.
arXiv:2608. 06839v1 Announce Type: new Abstract: Artificial Neural networks (ANNs) are often treated as black-box models, making explainability a central challenge in deep learning.
arXiv:2411. 18714v3 Announce Type: replace-cross Abstract: Self-driving cars increasingly rely on deep neural networks to achieve human-like driving.
arXiv:2608. 13456v1 Announce Type: new Abstract: World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution.
World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics.
arXiv:2607. 00267v1 Announce Type: cross Abstract: A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms.
arXiv:2605. 03413v2 Announce Type: replace Abstract: What does it mean to understand the world?
Structural causal models are the standard language for reasoning about interventions and counterfactuals, but they describe static variables, typically measured once, and usually forbid cyclic dependencies. Many systems we care about, such as patients, climates, and economies, instead evolve continuously in time, are observed at irregular time points, and contain feedback loops.
arXiv:2609.37680v1 Announce Type: cross Abstract: One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can...
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
arXiv:2607. 08641v1 Announce Type: new Abstract: Over the last few years, there has been an increased interest in making machine learning models more interpretable.