Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
3d ago

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.

By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud
arXiv AI
3d ago

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...

By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
arXiv AI
3d ago

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

The paper introduces the concept of "bottling"—the ability of large language model (LLM) agents to transform general capabilities into task‑specific, cost‑effective solutions for large, repetitive workloads. It presents BOTTLED, a benchmark where agents receive an unlabelled workload and must complete it within fixed time, compute, and API budgets, choosing strategies such as training small models or writing reusable programs. Experiments across ten models and three tasks show that strong zero‑shot performance does not guarantee effective bottling, yet bottling can still achieve substantial cost savings and competitive performance compared to specialized cheap inference models.

By Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh
arXiv AI
3d ago

APEX: Active Protection at Execution Boundaries for LLM Agents

The paper introduces APEX, an active defense for large language model agents that protects against indirect prompt injection by enforcing safety at execution boundaries. APEX uses an evidence‑gated prevention contract and deception‑based exposure to ensure that only authorized effects, endorsed by the task, are executed. Evaluation shows APEX achieves near‑zero attack success across multiple benchmarks and capability‑unit types, outperforming 13 baseline defenses.

By Xinran Zheng, Xin Fan Guo, Zhiqiang Hao, Fan Yang, Xingzhi Qian, Jiawei Du, Jinfeng Xu, Zheng Xing, Shuo Yang, Xingjun Wang