arXiv Machine Learning

PhysSAE: Mechanistic Interpretability of PINNs with Sparse Autoencoders

arXiv Machine Learning
Jul 15

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

arXiv:2607. 12166v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms.

By Mohamed Abdessalem Bal
arXiv AI
Jul 3

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

arXiv:2607. 01336v1 Announce Type: cross Abstract: Neural Quantum States (NQS) are a remarkably expressive class of variational ans\"atze for quantum many-body wavefunctions, yet little is understood about their internal mechanisms: trained on variational objectives alone, how do NQS accurately capture physical observables that they have never been explicitly optimized for?

By Zihao Qi, Christopher Earls
arXiv Machine Learning
Aug 27

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

The paper applies sparse autoencoders to a neutrino foundation model trained on IceCube data, uncovering a validated atlas of physical concepts within the model’s internal representation. Causal analysis shows the direction reconstruction head largely ignores this atlas, whereas an uncertainty head trained on the same representation effectively uses quality and brightness features, improving angular resolution from 20.2° to 3.2° at 20% efficiency. These findings demonstrate that mechanistic interpretability can expose latent physics and guide the design of downstream tasks.

By Rapha\"el Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri
Hugging Face Trending Papers
Jul 13

SPARC-Net: A Spectral, Causality-Aware, and Hard-Constrained Physics-Informed Architecture for Stiff and Shock-Dominated Partial Differential Equations

Physics-Informed Neural Networks (PINNs) provide a meshless approach for solving partial differential equations (PDEs), but suffer severe degradation in stiff and shock-dominated problems, where small PDE residuals can correspond to globally inaccurate solutions. We show these failures are multi-causal, arising from the concurrent interplay of (i) spectral bias against sharp features, (ii) imbalanced multi-term optimization and loss-weight collapse, (iii) violation of temporal causality, and (iv) under-resolved collocation.

Hugging Face Trending Papers
Jun 10

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

Generative AI emulators are increasingly used in scientific domains where we already have strong theory, benchmarks, and physical intuition. This raises a central evaluation and interpretability question: when a foundation-style model can reproduce known continuum dynamics, what internal mechanism supports that behavior, is the internal behaviour consistent with known physics, and how does it relate to where the emulator succeeds or fails?

arXiv Machine Learning
Jul 27

Neural Feature Governance: Extending Atom Prevalence

arXiv:2607. 21671v1 Announce Type: new Abstract: Neural network compression and interpretability remain open challenges in modern deep learn- ing, where billion-parameter architectures deliver impressive accuracy at the cost of trans- parency, computational efficiency, and reliable uncertainty quantification.

By Idris Karel Seunda Ekwe, Patrick Tenga Shako, Ernest Parfait Fokou\'e
arXiv Machine Learning
Jul 14

SPARC-Net: A Spectral, Causality-Aware, and Hard-Constrained Physics-Informed Architecture for Stiff and Shock-Dominated Partial Differential Equations

arXiv:2607. 11310v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) provide a meshless approach for solving partial differential equations (PDEs), but suffer severe degradation in stiff and shock-dominated problems, where small PDE residuals can correspond to globally inaccurate solutions.

By Divyavardhan Singh, Dimple Sonone, Hammad Mohammad, Kishor Upla
arXiv AI
1d ago

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.

By Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary