OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.
By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
The study investigates how large language models (LLMs) compensate for missing financial information by substituting user identity cues. Using 96,600 prompts to Llama‑3.1‑8B‑Instruct, the authors varied the amount of financial facts provided while keeping the underlying finances constant, and measured changes in recommended equity allocations across 100 financial profiles, 138 personas, and seven disclosure conditions. Results show that as financial facts are removed, the influence of identity on advice grows dramatically—from 5 % of variation with full disclosure to 96 % with none—while household size becomes the most reliable predictor when evidence is scarce, and gender effects persist even after controlling for standard errors.
"whyItMatters":"The findings highlight that LLM‑based advisory systems can produce biased financial recommendations when users provide incomplete information, underscoring the need for audits that reflect real‑world disclosure levels and consider the full spectrum of user identities."
By Saanvi Khetan, Sankar Balasubramanian
The paper introduces Agentic Semantic Sensing (Agentic SemS), a closed‑loop framework for AI‑enabled radio access networks that dynamically adjusts sensing configurations within a communication‑feasible profile set. A profile‑conditioned causal Transformer updates semantic beliefs from streaming data, while a semantic utility network guides the selection of the next sensing profile and determines when to exit early, balancing task benefit against sensing cost. Experiments on Widar3.0 demonstrate that Agentic SemS reduces cumulative sensing cost by 25.33% compared to a full‑sequence baseline while maintaining 85.79% Macro‑F1, and that semantic early exit yields an additional 12.35% cost savings with minimal performance loss.
By Zhongqin Wang, Xiaoqi Zhang, Nan Yang, Kai Wu, J. Andrew Zhang, Y. Jay Guo
The paper introduces DHCG, a framework that dynamically constructs hierarchical collaboration graphs for large language model–based multi‑agent systems. DHCG coordinates Planner, Worker, and Generator modules to adaptively determine the composition and scale of agents during execution, guided by feedback and action‑aware preference optimization. Experiments on code generation, mathematical reasoning, and domain‑specific tasks show DHCG surpasses static and dynamic baselines, improving performance by 2.77–8.02 points and achieving a 13.06‑point gain over single‑agent baselines.
By Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen
ShanLiangRen is a nutrition agent designed to generate personalized daily meal plans that satisfy both user constraints and multidimensional nutritional goals. It transforms dietary specifications, nutrient data, user attributes, and natural language requirements into a constrained planning instance, then uses a retrieval‑augmented generation approach to narrow the candidate set from a large ingredient and recipe space. Finally, it refines plans via Pareto‑guided iterative revisions with an LLM, producing fully quantified meal plans with explicit ingredients, portion sizes, and compliance reports.
By Miao Xie, Xiao Zhang, Yuan Wang, Ruixin Zhu, Chunli Lv
The paper introduces a method for sea surface temperature (SST) forecasting that combines textual environmental context with spatial graph representations for large language models (LLMs). Historical SST and anomaly sequences, date‑aligned environmental records, and static ocean knowledge are provided as textual input, while a static graph captures geographic–climatological relations and a dynamic graph captures recent SST correlations and tropical‑cyclone influence. The approach achieves the lowest mean absolute error and highest R² among compared methods over ten forecast steps in the South China Sea, and includes a rule‑based module that links predicted trends to source‑linked contextual explanations.
By Xiong Li, Xiaowei Zhou, Yanwei Yu, Qian Cui, Junyu Dong
The paper introduces Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an LLM agent has successfully completed its task using a single trajectory without requiring privileged model access or training data. CRGs decompose the agent’s overall claim of success into contextualized sub-claims, estimate confidence for each terminal claim based on trajectory evidence, and aggregate these into an overall confidence estimate. Experiments across multiple benchmarks, models, and agent frameworks show that CRGs produce better-calibrated confidence and more effective risk-aware decision making than existing verbalized, sampling-based, and white-box surrogate methods, while also providing transparent, auditable evidence for each estimate.
By Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka
DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.
By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud
arXiv:2610.08106v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...
By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
Learn2Play Bench is a new benchmark that tests how well large language model agents learn from experience in unfamiliar, text‑based games with novel or counterintuitive rules. The benchmark provides reproducible feedback, automatic scoring, and varied game instances to evaluate learning across repeated attempts and transfer to new situations. Findings show that retaining full action records aids learning, human players outperform agents, and the choice of harness significantly impacts performance and inference cost.
By Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi
arXiv:2610.08216v1 Announce Type: new
Abstract: Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pool...
By Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu
arXiv:2610.08246v1 Announce Type: new
Abstract: Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing plannin...
By Andr\'{e} G. Pereira, Augusto B. Corr\^ea, Felipe Meneguzzi, Jendrik Seipp
arXiv:2610.08250v1 Announce Type: new
Abstract: Large language models are increasingly used to simulate clients for counselor training and psychological counseling research, but reliable simulation r...
By Shixin Peng, Kun Jiang, Jiaxing Zheng, Qihao Yang, Jingying Chen
arXiv:2610.08312v1 Announce Type: new
Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, rec...
By Maoqi Liu, Quan Fang, Yufei He
arXiv:2610.08319v1 Announce Type: new
Abstract: In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails...
By Hanchao Zhou, Jialei Li
arXiv:2610.08446v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physi...
By Zhiyuan Qi, Jierui Li, Yifan Shen, Cheng Qian, Jiateng Liu
arXiv:2610.08586v1 Announce Type: new
Abstract: Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complet...
By Sujato Dutta, Sreekruthy Tummala, Shashank Vanga, Ayushmi Pavani
arXiv:2610.08647v1 Announce Type: new
Abstract: LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple...
By Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu
arXiv:2610.08720v1 Announce Type: new
Abstract: LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical...
By Siru Jiang, Yongzhe Lyu, Shuo Lu, Yubin Wang, Yuxiang Zhang, Yue Liao, Bin Wang, Jian Liang, Tieniu Tan
The paper introduces the concept of "bottling"—the ability of large language model (LLM) agents to transform general capabilities into task‑specific, cost‑effective solutions for large, repetitive workloads. It presents BOTTLED, a benchmark where agents receive an unlabelled workload and must complete it within fixed time, compute, and API budgets, choosing strategies such as training small models or writing reusable programs. Experiments across ten models and three tasks show that strong zero‑shot performance does not guarantee effective bottling, yet bottling can still achieve substantial cost savings and competitive performance compared to specialized cheap inference models.
By Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh