A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning
arXiv:2601. 09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models.
arXiv:2606. 16920v1 Announce Type: cross Abstract: Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task.
arXiv:2601. 09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models.
arXiv:2607. 18921v1 Announce Type: cross Abstract: Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior.
arXiv:2601. 21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive.
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
arXiv:2605. 28860v2 Announce Type: replace-cross Abstract: Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities.
arXiv:2606. 09899v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to identify which internal components causally drive a language model's behavior.
arXiv:2606. 24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize.
arXiv:2510. 12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes.
arXiv:2605. 29548v2 Announce Type: replace Abstract: Larger models learn tasks smaller models do not.
arXiv:2605. 24059v2 Announce Type: replace Abstract: We present a three-step recipe for identifying attention-head circuits in pretrained transformers.
arXiv:2608. 05224v1 Announce Type: new Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions.
arXiv:2608. 09117v1 Announce Type: new Abstract: Probabilistic Circuits (PCs) are tractable generative models whose internal nodes encode a hierarchy of probabilistic sum- maries over different variable scopes.