arXiv AI

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

arXiv:2607. 16448v1 Announce Type: cross Abstract: Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear probes, PCA, SVD) to more expensive training-based ones (such as SAEs and activation oracles).

arXiv AI
Sep 10

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.

By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing
arXiv AI
Jul 9

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

By Kabir Moghe, Peter Chin
arXiv AI
Aug 20

What is Missing from AI Post-Training AI: An Empirical Analysis

The paper investigates the limitations of post-training AI agents that can autonomously train large language models. It distinguishes between execution-level capability—making adjustments within a chosen training strategy—and strategy-level capability—revising the overall approach based on new evidence. Analysis of many public post-training runs shows that agents lock into a strategy early and then only perform local tweaks, regardless of task. Experiments with experience scaffolds, human guidance, and extra compute improve execution but do not enable strategy reevaluation, indicating that agents lack a mechanism to spontaneously reassess their strategy during training.

By Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu, Xin Cong, Zhong Zhang, Maosong Sun, Yankai Lin
arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow
arXiv Machine Learning
Sep 22

Characterizing Model-Native Skills

arXiv:2604.17614v2 Announce Type: replace-cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterization...

By Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia