Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv AI
1d ago

Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

The paper evaluates the impact of adding a persistent memory tier to a multi‑agent large language model inference system. While decomposing long‑context inference across agents reduces the peak KV cache usage from about 35 MiB to 14.3 MiB per query, the persistent tier adds only a modest 0.368 MiB to the cache and shows no measurable accuracy improvement across eight dataset pairs. The authors argue that the lack of benefit is structural, as single‑question benchmarks do not provide informative recall opportunities, and they outline conditions and detection procedures for agent‑memory ablations.

By Hochan Son, Kyungdoe Han, Jaehan Koh, Xiaowu Dai, Wenlu Xu, Guang Cheng
arXiv AI
1d ago

OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation

OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.

By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
arXiv AI
1d ago

DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM-Based Multi-Agent Reasoning

The paper introduces DHCG, a framework that dynamically constructs hierarchical collaboration graphs for large language model–based multi‑agent systems. DHCG coordinates Planner, Worker, and Generator modules to adaptively determine the composition and scale of agents during execution, guided by feedback and action‑aware preference optimization. Experiments on code generation, mathematical reasoning, and domain‑specific tasks show DHCG surpasses static and dynamic baselines, improving performance by 2.77–8.02 points and achieving a 13.06‑point gain over single‑agent baselines.

By Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen
arXiv AI
1d ago

RadOnc-Agent: An LLM-Orchestrated Framework for AI Workflows Across the Radiotherapy Care Pathway

RadOnc-Agent is an AI framework that organizes radiotherapy into four clinical phases and offers 26 callable functions via a conversational interface. A large‑language‑model controller maps clinical intent to schema‑constrained calls, maintains patient and workflow context, and routes requests to specialist services. In evaluations, the system achieved high accuracy in single‑function calls (98.79%) and cross‑stage workflow completions (96.50% scripted, 96.67% real‑patient), demonstrating technical feasibility of LLM‑orchestrated coordination across heterogeneous radiotherapy capabilities.

By Caiwen Jiang, Shuoyang Wei, Songlin Zhao, Junyu Li, Jingyuan Chen, Wei Liu
arXiv AI
1d ago

Textual Environmental Context and Spatial Graphs for LLM-Based Regional SST Forecasting

The paper introduces a method for sea surface temperature (SST) forecasting that combines textual environmental context with spatial graph representations for large language models (LLMs). Historical SST and anomaly sequences, date‑aligned environmental records, and static ocean knowledge are provided as textual input, while a static graph captures geographic–climatological relations and a dynamic graph captures recent SST correlations and tropical‑cyclone influence. The approach achieves the lowest mean absolute error and highest R² among compared methods over ten forecast steps in the South China Sea, and includes a rule‑based module that links predicted trends to source‑linked contextual explanations.

By Xiong Li, Xiaowei Zhou, Yanwei Yu, Qian Cui, Junyu Dong
arXiv AI
1d ago

Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy

The paper presents a defense‑in‑depth framework for large language models that separates internal activation steering from external memory handling to combat sycophancy. It evaluates four open‑weight models on a new MemSyco‑Bench dataset, testing five memory‑defense configurations—including a Router Gate that selectively rewrites, keeps, or drops memories—and measures sycophancy and accuracy across 1,550 items. Results show that selective Router Gate filtering preserves more accuracy than complete memory removal, while inverse steering slightly reduces sycophancy but is not statistically significant.

By Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary
arXiv AI
1d ago

Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs

The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.

By Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan
arXiv AI
1d ago

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.

By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
arXiv AI
1d ago

Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals

The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
arXiv Machine Learning
1d ago

Skillful Data-Driven Subseasonal Soil Moisture Forecasting: Prospects and Limits for Flash Drought Prediction

The paper presents a Vision Transformer-based model for subseasonal soil‑moisture forecasting over Europe, showing that forecast skill depends heavily on how the prediction problem is formulated. By using residual learning and forecasting root‑zone soil moisture in physical units, the model outperforms persistence and existing deep‑learning and ECMWF baselines, providing well‑calibrated probabilistic predictions. However, predicting flash drought onset—defined by rapid multi‑pentad intensification—remains a challenge shared by all current subseasonal‑to‑seasonal systems.

By Noelia Otero, Atahan \"Ozer, Miguel-\'Angel Fern\'andez-Torres, Jackie Ma
arXiv Machine Learning
1d ago

Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints

The paper introduces HARP, a hard-budget allocation method for evaluating large language models (LLMs) in multi-turn interactions where the time-to-event is partially observed due to resource limits. HARP guarantees that the total computational budget is never exceeded, reallocates unused budget, and provides lower predictive bounds (LPBs) with finite-sample coverage and unbiased metric estimates. Experiments on tasks such as jailbreaks, toxic content, and hallucinations demonstrate that HARP achieves near-nominal coverage with low variance while respecting the fixed budget.

By Shai Feldman, Yaniv Romano
arXiv Machine Learning
1d ago

Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes

The paper investigates how exporting ternary language models (BitNet, Falcon‑E, BitCPM) through a bf16 cast step can introduce significant discrepancies between the fine‑tuned latent weights and the deployed ternary codes. In three lab pipelines, the authors find that fp32 quantization of shipped latents disagrees with the deployed codes on up to 1.77% of codes, and that the export step can drastically reduce strict accuracy on GSM8K (e.g., from 58.79% to 0.78% for Falcon‑E‑1B‑Base). They propose two compatibility remedies—directly writing the training quantizer’s codes or adjusting bf16 inputs—to meet a 4‑point strict‑accuracy non‑inferiority criterion across all models.

By Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness)
arXiv AI
1d ago

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

arXiv:2610.08106v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...

By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
arXiv Machine Learning
1d ago

ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding

ProximalFM is a transformer‑based model that uses prior‑data fitted networks (PFNs) to perform Bayesian proximal causal inference under hidden confounding. By training on synthetic data generated from structural causal models with oracle counterfactuals, it amortizes the Bayesian operator inversion into a single forward pass, producing posterior estimates of the conditional average treatment effect (CATE). The approach consistently outperforms prior methods across various proximal regimes, especially when latent confounding is strong and proxy variables are weakly informative, and it requires no dataset‑specific tuning.

By Christophe Muller, Ayub Kharel, Alex Luedtke, Chan Park, Eric Tchetgen Tchetgen, Juan L. Gamella, Rahul Krishnan, Ricardo Silva, Jakob Zeitler
arXiv Machine Learning
1d ago

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models

The paper introduces Mask Fine‑Tuning (MFT), a new approach for adapting Vision‑Language Models that avoids modifying backbone weights. MFT learns masks to selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align with downstream tasks. Experiments demonstrate that MFT consistently outperforms both Full Fine‑Tuning and Parameter‑Efficient Fine‑Tuning across multiple benchmarks, while also offering insights into how pretrained VLMs reorganize their internal pathways during adaptation.

By Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu
arXiv Computation and Language
1d ago

Structured but Silent: Probing Capability Requirements in LLM Hidden States

The paper investigates whether large language models (LLMs) can infer the capability requirements of a user query before generating a response. Using the TACIT framework, the authors decompose these requirements into eight classes across three axes—Source, Transformation, and World Effect—and train linear probes on hidden states from four open-weight LLM families. Their results show that these capability structures are linearly decodable with high accuracy, yet the models struggle to express the same information in natural language, a phenomenon termed "structured but silent."

By Kyojun Choo, Minsoo Song, Yunju Kang, Chanjun Park
arXiv Computation and Language
1d ago

The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception

The paper introduces EmoNet‑Face‑HQ, a fine‑grained emotion recognition benchmark that uses generated portraits and a 40‑category taxonomy to evaluate vision‑language models (VLMs). It finds that VLMs perform poorly when asked to generate responses but can match or surpass a fine‑tuned model (Empathic‑Insight‑Face) when their logits are read directly as binary queries. The study shows that the benchmark’s difficulty lies in the readout process rather than in perception, and that graded probability outputs yield better performance than simple yes/no questions.

By Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth Andr\'e
arXiv Computer Vision
1d ago

WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.

By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv Computer Vision
1d ago

Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering

The paper proposes a production‑centered framework for comparing detection and watermarking techniques across two main AI image and video creation routes: direct visual generation and LLM‑driven code rendering. It introduces an explicit verification specification that separates passive inference, message recovery, and authenticated provenance, and organizes watermarks by production stage for images, videos, source code, and rendering‑aware outputs. The authors outline ten research questions covering identifiability, observability, fair comparison, payload recoverability, reconstruction, synchronization, composition, hybrid local contribution, and private production‑event authentication, and connect the framework to concrete systems such as Claude, OpenAI, and rendering tools.

By Zheng Gao, Xiaoyu Li, Zhicheng Bao, Yang Song, Jiaojiao Jiang