arXiv AI

Agentic Neural Architecture Search

arXiv:2607. 07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task.

arXiv AI
Jun 9

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

arXiv:2606. 09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite.

By Pu Ning, Quan Chen, Kun Tao, Xinyu Tang, Tianshu Wang, Qianggang Cao, Xinyu Kong, Zujie Wen, Zhiqiang Zhang, Jun Zhou
Hugging Face Trending Papers
Aug 3

Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study

We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records experiments over an extended period.

arXiv Machine Learning
Sep 25

EvoTreeNAD: Genealogy-Guided Evolution for LLM-Driven Neural Architecture Discovery

EvoTreeNAD is a genealogy‑guided evolutionary algorithm that autonomously discovers neural architectures without a predefined seed or search space. Starting from an empty root, it builds a persistent genealogy where each node represents a complete architecture; top‑percentile values from nodes and descendants steer lineage selection. The method combines an Idea Agent that proposes variants and a Code Agent that implements them, with theoretical analysis showing stationary variation regimes and empirical results demonstrating superior performance on CIFAR‑10/100 and MedMNIST‑v2 tasks.

By Lishan Yu, Derek Jiu, Qizhen Lan, Xiaoqian Jiang
arXiv AI
Jul 9

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

By Kabir Moghe, Peter Chin
arXiv AI
Aug 25

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.

By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
arXiv AI
Jul 10

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

arXiv:2607. 08662v1 Announce Type: cross Abstract: Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex, deep-and-wide search and research-oriented tasks.

By Xiaoshuai Song, Liancheng Zhang, Kangzhi Zhao, Yutao Zhu, Zhongyuan Wang, Guanting Dong, Jinghan Yang, Han Li, Kun Gai, Ji-Rong Wen, Zhicheng Dou
arXiv AI
3d ago

Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

The paper introduces a self‑evolving harness framework where a frozen language‑model agent first solves tasks and then edits its own harness based on run records. Using a 49‑line seed harness, the evolved harness improves average scores on in‑distribution benchmarks by 4.48 points and on out‑of‑distribution benchmarks by 12.64 points, surpassing Codex on the former and matching it on the latter. Continued evolution on a specific out‑of‑distribution benchmark further raises performance, and the study analyzes emergent mechanisms such as output truncation and history compaction.

By Qiankai Xu