arXiv Machine Learning

MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection

MECHVAR is a lightweight, auditable rule for selecting experiments from a finite library to discriminate between candidate mechanisms. It chooses probes by maximizing the posterior‑weighted variance of predicted responses, a score that aligns with the Box–Hill pairwise‑KL criterion and links to expected information gain when separations are small. Experiments on a 25‑block audit and a Digits loop show MECHVAR outperforming confirmation‑first strategies and matching or exceeding EIG in identification accuracy while being far faster to compute.

arXiv AI
Sep 10

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.

By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
arXiv Machine Learning
Sep 24

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.

By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang
arXiv Machine Learning
Aug 4

Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs

arXiv:2608. 01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits.

By Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
arXiv Machine Learning
1d ago

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.

By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
arXiv Machine Learning
Sep 11

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients

The paper proposes a flexible four‑coefficient parameterization for per‑token gating in on‑policy knowledge distillation, unifying existing methods such as EOPD and ToDi as special cases. Experiments on TweetEval with a Qwen3 teacher‑student pair show that the full family of gating configurations outperforms single‑channel baselines in most cells, and dynamic gating beats static baselines in a majority of isolated comparisons. The authors present the framework mainly as a shared coordinate system for comparing gating designs rather than as definitive statistical evidence.

By Suwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang