arXiv AI By Vijay Erramilli

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Read the original on arXiv AI →

arXiv:2608. 19338v1 Announce Type: cross Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 12

Order Is Not Control

arXiv:2606. 12923v1 Announce Type: cross Abstract: AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects.

By Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Tim Elson