DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.
By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud
arXiv:2610.08106v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...
By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
arXiv:2610.08216v1 Announce Type: new
Abstract: Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pool...
By Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu
arXiv:2610.08246v1 Announce Type: new
Abstract: Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing plannin...
By Andr\'{e} G. Pereira, Augusto B. Corr\^ea, Felipe Meneguzzi, Jendrik Seipp
arXiv:2610.08250v1 Announce Type: new
Abstract: Large language models are increasingly used to simulate clients for counselor training and psychological counseling research, but reliable simulation r...
By Shixin Peng, Kun Jiang, Jiaxing Zheng, Qihao Yang, Jingying Chen
arXiv:2610.08312v1 Announce Type: new
Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, rec...
By Maoqi Liu, Quan Fang, Yufei He
arXiv:2610.08319v1 Announce Type: new
Abstract: In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails...
By Hanchao Zhou, Jialei Li
arXiv:2610.08446v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physi...
By Zhiyuan Qi, Jierui Li, Yifan Shen, Cheng Qian, Jiateng Liu
arXiv:2610.08586v1 Announce Type: new
Abstract: Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complet...
By Sujato Dutta, Sreekruthy Tummala, Shashank Vanga, Ayushmi Pavani
arXiv:2610.08647v1 Announce Type: new
Abstract: LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple...
By Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu
arXiv:2610.08720v1 Announce Type: new
Abstract: LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical...
By Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan
The paper introduces the concept of "bottling"—the ability of large language model (LLM) agents to transform general capabilities into task‑specific, cost‑effective solutions for large, repetitive workloads. It presents BOTTLED, a benchmark where agents receive an unlabelled workload and must complete it within fixed time, compute, and API budgets, choosing strategies such as training small models or writing reusable programs. Experiments across ten models and three tasks show that strong zero‑shot performance does not guarantee effective bottling, yet bottling can still achieve substantial cost savings and competitive performance compared to specialized cheap inference models.
By Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh
arXiv:2610.08778v1 Announce Type: new
Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach...
By Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang
arXiv:2610.05094v1 Announce Type: cross
Abstract: Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design...
By Laur\`ene Vaugrante, Thilo Hagendorff
arXiv:2610.06892v1 Announce Type: cross
Abstract: Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are l...
By Soumya Nasipuri, Sayak Ray Chowdhury, Sanjukta Roy
arXiv:2610.06927v1 Announce Type: cross
Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free r...
By Sara Abdali, Jongwoo Ko, Pashmina Cameron
The paper introduces APEX, an active defense for large language model agents that protects against indirect prompt injection by enforcing safety at execution boundaries. APEX uses an evidence‑gated prevention contract and deception‑based exposure to ensure that only authorized effects, endorsed by the task, are executed. Evaluation shows APEX achieves near‑zero attack success across multiple benchmarks and capability‑unit types, outperforming 13 baseline defenses.
By Xinran Zheng, Xin Fan Guo, Zhiqiang Hao, Fan Yang, Xingzhi Qian, Jiawei Du, Jinfeng Xu, Zheng Xing, Shuo Yang, Xingjun Wang
arXiv:2610.06977v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-sou...
By Xiaojun Jia, Simeng Qin, Yiming Li, Jie Liao, Sensen Gao, Ke Ma, Yang Liu, Xiaochun Cao
arXiv:2610.06993v1 Announce Type: cross
Abstract: Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However,...
By Zhishen Sun, Hongzhan Wang, Sizhe Dang, Guang Dai, Haishan Ye
arXiv:2610.06996v1 Announce Type: cross
Abstract: Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memor...
By Gleb Molodtsov, Ekaterina Alimaskina, Evgeny Uskov, Artur Zagitov, Aleksandr Beznosikov