arXiv Machine Learning By Yifan Guo

MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection

Read the original on arXiv Machine Learning →

MECHVAR is a lightweight, auditable rule for selecting experiments from a finite library to discriminate between candidate mechanisms. It chooses probes by maximizing the posterior‑weighted variance of predicted responses, a score that aligns with the Box–Hill pairwise‑KL criterion and links to expected information gain when separations are small. Experiments on a 25‑block audit and a Digits loop show MECHVAR outperforming confirmation‑first strategies and matching or exceeding EIG in identification accuracy while being far faster to compute.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.

By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
arXiv Machine Learning
Sep 24

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.

By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang