Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv AI
2d ago

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.

By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
arXiv AI
2d ago

Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts

The study investigates how large language models (LLMs) compensate for missing financial information by substituting user identity cues. Using 96,600 prompts to Llama‑3.1‑8B‑Instruct, the authors varied the amount of financial facts provided while keeping the underlying finances constant, and measured changes in recommended equity allocations across 100 financial profiles, 138 personas, and seven disclosure conditions. Results show that as financial facts are removed, the influence of identity on advice grows dramatically—from 5 % of variation with full disclosure to 96 % with none—while household size becomes the most reliable predictor when evidence is scarce, and gender effects persist even after controlling for standard errors. "whyItMatters":"The findings highlight that LLM‑based advisory systems can produce biased financial recommendations when users provide incomplete information, underscoring the need for audits that reflect real‑world disclosure levels and consider the full spectrum of user identities."

By Saanvi Khetan, Sankar Balasubramanian
arXiv AI
2d ago

Agentic Semantic Sensing for Resource-Adaptive AI-RAN

The paper introduces Agentic Semantic Sensing (Agentic SemS), a closed‑loop framework for AI‑enabled radio access networks that dynamically adjusts sensing configurations within a communication‑feasible profile set. A profile‑conditioned causal Transformer updates semantic beliefs from streaming data, while a semantic utility network guides the selection of the next sensing profile and determines when to exit early, balancing task benefit against sensing cost. Experiments on Widar3.0 demonstrate that Agentic SemS reduces cumulative sensing cost by 25.33% compared to a full‑sequence baseline while maintaining 85.79% Macro‑F1, and that semantic early exit yields an additional 12.35% cost savings with minimal performance loss.

By Zhongqin Wang, Xiaoqi Zhang, Nan Yang, Kai Wu, J. Andrew Zhang, Y. Jay Guo
arXiv AI
2d ago

DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM-Based Multi-Agent Reasoning

The paper introduces DHCG, a framework that dynamically constructs hierarchical collaboration graphs for large language model–based multi‑agent systems. DHCG coordinates Planner, Worker, and Generator modules to adaptively determine the composition and scale of agents during execution, guided by feedback and action‑aware preference optimization. Experiments on code generation, mathematical reasoning, and domain‑specific tasks show DHCG surpasses static and dynamic baselines, improving performance by 2.77–8.02 points and achieving a 13.06‑point gain over single‑agent baselines.

By Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen
arXiv AI
2d ago

ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning

ShanLiangRen is a nutrition agent designed to generate personalized daily meal plans that satisfy both user constraints and multidimensional nutritional goals. It transforms dietary specifications, nutrient data, user attributes, and natural language requirements into a constrained planning instance, then uses a retrieval‑augmented generation approach to narrow the candidate set from a large ingredient and recipe space. Finally, it refines plans via Pareto‑guided iterative revisions with an LLM, producing fully quantified meal plans with explicit ingredients, portion sizes, and compliance reports.

By Miao Xie, Xiao Zhang, Yuan Wang, Ruixin Zhu, Chunli Lv
arXiv AI
2d ago

Textual Environmental Context and Spatial Graphs for LLM-Based Regional SST Forecasting

The paper introduces a method for sea surface temperature (SST) forecasting that combines textual environmental context with spatial graph representations for large language models (LLMs). Historical SST and anomaly sequences, date‑aligned environmental records, and static ocean knowledge are provided as textual input, while a static graph captures geographic–climatological relations and a dynamic graph captures recent SST correlations and tropical‑cyclone influence. The approach achieves the lowest mean absolute error and highest R² among compared methods over ten forecast steps in the South China Sea, and includes a rule‑based module that links predicted trends to source‑linked contextual explanations.

By Xiong Li, Xiaowei Zhou, Yanwei Yu, Qian Cui, Junyu Dong
arXiv AI
2d ago

Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

The paper introduces Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an LLM agent has successfully completed its task using a single trajectory without requiring privileged model access or training data. CRGs decompose the agent’s overall claim of success into contextualized sub-claims, estimate confidence for each terminal claim based on trajectory evidence, and aggregate these into an overall confidence estimate. Experiments across multiple benchmarks, models, and agent frameworks show that CRGs produce better-calibrated confidence and more effective risk-aware decision making than existing verbalized, sampling-based, and white-box surrogate methods, while also providing transparent, auditable evidence for each estimate.

By Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka
arXiv AI
2d ago

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.

By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud
arXiv AI
2d ago

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...

By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
arXiv AI
2d ago

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Learn2Play Bench is a new benchmark that tests how well large language model agents learn from experience in unfamiliar, text‑based games with novel or counterintuitive rules. The benchmark provides reproducible feedback, automatic scoring, and varied game instances to evaluate learning across repeated attempts and transfer to new situations. Findings show that retaining full action records aids learning, human players outperform agents, and the choice of harness significantly impacts performance and inference cost.

By Yibo Li, Jinhang Qiu, Zhi Zheng, Qianyun Guo, Jiaying Wu, Shuo Ji, Bryan Hooi
arXiv AI
2d ago

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

The paper introduces the concept of "bottling"—the ability of large language model (LLM) agents to transform general capabilities into task‑specific, cost‑effective solutions for large, repetitive workloads. It presents BOTTLED, a benchmark where agents receive an unlabelled workload and must complete it within fixed time, compute, and API budgets, choosing strategies such as training small models or writing reusable programs. Experiments across ten models and three tasks show that strong zero‑shot performance does not guarantee effective bottling, yet bottling can still achieve substantial cost savings and competitive performance compared to specialized cheap inference models.

By Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh