arXiv Machine Learning

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

arXiv:2608. 05207v1 Announce Type: new Abstract: Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning.

arXiv AI
Aug 5

Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators

arXiv:2608. 02712v1 Announce Type: cross Abstract: Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness through pipelines that couple LLMs with agentic reinforcement learning and evolutionary search.

By Yansong Sun, Shenxiu Wu, Siyuan Chen, Runlin Hou, Junhao Qiu, Junming Cao, Shudi Shao, Zhichao Lu, Qingfu Zhang
arXiv AI
Jul 9

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

By Kabir Moghe, Peter Chin
arXiv AI
Jun 18

SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents

arXiv:2606. 18733v1 Announce Type: cross Abstract: Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generation, or benchmark-driven model selection.

By Qiao Zhao, JianYing Qu, Jun Zhang, Yehua Yang, Hanwen Du, Zhongkai Sun
arXiv AI
Sep 10

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.

By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu