The paper evaluates the impact of adding a persistent memory tier to a multi‑agent large language model inference system. While decomposing long‑context inference across agents reduces the peak KV cache usage from about 35 MiB to 14.3 MiB per query, the persistent tier adds only a modest 0.368 MiB to the cache and shows no measurable accuracy improvement across eight dataset pairs. The authors argue that the lack of benefit is structural, as single‑question benchmarks do not provide informative recall opportunities, and they outline conditions and detection procedures for agent‑memory ablations.
By Hochan Son, Kyungdoe Han, Jaehan Koh, Xiaowu Dai, Wenlu Xu, Guang Cheng
OOPMAS introduces a training‑free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are defined as object‑oriented class definitions with dedicated roles, tools, and persistent state, while workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in‑context improvement without any gradient updates or fine‑tuning, and achieves 89.6% accuracy on a mixed‑task benchmark, outperforming the strongest baseline by 18.1 percentage points.
By Qi Cheng, Shengyu Chen, Wei Cheng, Yiqun Xie, Xiaowei Jia, Haoyu Wang, Haifeng Chen
The paper introduces DHCG, a framework that dynamically constructs hierarchical collaboration graphs for large language model–based multi‑agent systems. DHCG coordinates Planner, Worker, and Generator modules to adaptively determine the composition and scale of agents during execution, guided by feedback and action‑aware preference optimization. Experiments on code generation, mathematical reasoning, and domain‑specific tasks show DHCG surpasses static and dynamic baselines, improving performance by 2.77–8.02 points and achieving a 13.06‑point gain over single‑agent baselines.
By Jie Ren, Jiakang Yuan, Chenyu Huang, Hezeer Ma, Jiayuan Fan, Tao Chen
RadOnc-Agent is an AI framework that organizes radiotherapy into four clinical phases and offers 26 callable functions via a conversational interface. A large‑language‑model controller maps clinical intent to schema‑constrained calls, maintains patient and workflow context, and routes requests to specialist services. In evaluations, the system achieved high accuracy in single‑function calls (98.79%) and cross‑stage workflow completions (96.50% scripted, 96.67% real‑patient), demonstrating technical feasibility of LLM‑orchestrated coordination across heterogeneous radiotherapy capabilities.
By Caiwen Jiang, Shuoyang Wei, Songlin Zhao, Junyu Li, Jingyuan Chen, Wei Liu
The paper introduces a method for sea surface temperature (SST) forecasting that combines textual environmental context with spatial graph representations for large language models (LLMs). Historical SST and anomaly sequences, date‑aligned environmental records, and static ocean knowledge are provided as textual input, while a static graph captures geographic–climatological relations and a dynamic graph captures recent SST correlations and tropical‑cyclone influence. The approach achieves the lowest mean absolute error and highest R² among compared methods over ten forecast steps in the South China Sea, and includes a rule‑based module that links predicted trends to source‑linked contextual explanations.
By Xiong Li, Xiaowei Zhou, Yanwei Yu, Qian Cui, Junyu Dong
The paper presents a defense‑in‑depth framework for large language models that separates internal activation steering from external memory handling to combat sycophancy. It evaluates four open‑weight models on a new MemSyco‑Bench dataset, testing five memory‑defense configurations—including a Router Gate that selectively rewrites, keeps, or drops memories—and measures sycophancy and accuracy across 1,550 items. Results show that selective Router Gate filtering preserves more accuracy than complete memory removal, while inverse steering slightly reduces sycophancy but is not statistically significant.
By Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary
The paper introduces TraceDSE, an agentic design space exploration framework for jointly mapping AI inference workloads to heterogeneous edge SoCs and configuring each processing unit. Unlike traditional black-box optimization, TraceDSE uses a proposer‑critic loop powered by large language models and enriched with system execution traces to identify bottlenecks and refine design choices. Experiments on an Intel Meteor Lake SoC show that TraceDSE outperforms state‑of‑the‑art evolutionary and Bayesian methods, improving Pareto frontier hypervolume by up to 68% while reducing hardware evaluations by 6–9×.
By Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan
VisionWeave introduces elastic visual representation weaving, a native capability for multimodal large language models that learns where and at what granularity to encode visual information. The method combines a gated spatial pooler for coarse representations with a granularity router that allocates content‑adaptive token usage, trained end‑to‑end on large‑scale data. Experiments on Qwen3.5‑4B and Qwen3.8‑27B show that VisionWeave can save 43.0% of tokens while preserving 98.9% of performance across eight benchmarks, and delivers significant throughput gains and latency reductions when deployed on the SGLang serving engine.
By Yuan Feng, Qize Yang, Ruizhe Chen, Sibo Song, Haolin He, Muzhi Zhu, Zihan Liu, Yunfei Chu, Xize Cheng, Yuxuan Wang, Jin Xu, Xike Xie
The paper presents a new multi‑turn benchmark of 423 conversations with 1,661 labeled turn‑states to study when language models should refrain from answering. It shows that probes for unanswerability transfer well across datasets that share the same underlying signal, but fail to generalise to other forms of epistemic uncertainty. While a calibrated probe can identify underspecified turns more accurately than chance, it does not consistently improve overall generation quality compared to standard methods.
By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Danil Fedorov, Kirill Redko, Sergey Chuprin, Aidar Shumbalov, Stanislav Chumakov, Anna Kalyuzhnaya
The paper presents a Vision Transformer-based model for subseasonal soil‑moisture forecasting over Europe, showing that forecast skill depends heavily on how the prediction problem is formulated. By using residual learning and forecasting root‑zone soil moisture in physical units, the model outperforms persistence and existing deep‑learning and ECMWF baselines, providing well‑calibrated probabilistic predictions. However, predicting flash drought onset—defined by rapid multi‑pentad intensification—remains a challenge shared by all current subseasonal‑to‑seasonal systems.
By Noelia Otero, Atahan \"Ozer, Miguel-\'Angel Fern\'andez-Torres, Jackie Ma
The paper introduces HARP, a hard-budget allocation method for evaluating large language models (LLMs) in multi-turn interactions where the time-to-event is partially observed due to resource limits. HARP guarantees that the total computational budget is never exceeded, reallocates unused budget, and provides lower predictive bounds (LPBs) with finite-sample coverage and unbiased metric estimates. Experiments on tasks such as jailbreaks, toxic content, and hallucinations demonstrate that HARP achieves near-nominal coverage with low variance while respecting the fixed budget.
By Shai Feldman, Yaniv Romano
The paper investigates how exporting ternary language models (BitNet, Falcon‑E, BitCPM) through a bf16 cast step can introduce significant discrepancies between the fine‑tuned latent weights and the deployed ternary codes. In three lab pipelines, the authors find that fp32 quantization of shipped latents disagrees with the deployed codes on up to 1.77% of codes, and that the export step can drastically reduce strict accuracy on GSM8K (e.g., from 58.79% to 0.78% for Falcon‑E‑1B‑Base). They propose two compatibility remedies—directly writing the training quantizer’s codes or adjusting bf16 inputs—to meet a 4‑point strict‑accuracy non‑inferiority criterion across all models.
By Avichal Sahai (Ofbusiness), Nishant Raj (Ofbusiness), Animesh Srivastava (Ofbusiness)
arXiv:2610.08106v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observe...
By Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen
arXiv:2610.08778v1 Announce Type: new
Abstract: Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach...
By Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang
ProximalFM is a transformer‑based model that uses prior‑data fitted networks (PFNs) to perform Bayesian proximal causal inference under hidden confounding. By training on synthetic data generated from structural causal models with oracle counterfactuals, it amortizes the Bayesian operator inversion into a single forward pass, producing posterior estimates of the conditional average treatment effect (CATE). The approach consistently outperforms prior methods across various proximal regimes, especially when latent confounding is strong and proxy variables are weakly informative, and it requires no dataset‑specific tuning.
By Christophe Muller, Ayub Kharel, Alex Luedtke, Chan Park, Eric Tchetgen Tchetgen, Juan L. Gamella, Rahul Krishnan, Ricardo Silva, Jakob Zeitler
The paper introduces Mask Fine‑Tuning (MFT), a new approach for adapting Vision‑Language Models that avoids modifying backbone weights. MFT learns masks to selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align with downstream tasks. Experiments demonstrate that MFT consistently outperforms both Full Fine‑Tuning and Parameter‑Efficient Fine‑Tuning across multiple benchmarks, while also offering insights into how pretrained VLMs reorganize their internal pathways during adaptation.
By Mingyuan Zhang, Yue Bai, Yifan Wang, Yiyang Huang, Yun Fu
The paper investigates whether large language models (LLMs) can infer the capability requirements of a user query before generating a response. Using the TACIT framework, the authors decompose these requirements into eight classes across three axes—Source, Transformation, and World Effect—and train linear probes on hidden states from four open-weight LLM families. Their results show that these capability structures are linearly decodable with high accuracy, yet the models struggle to express the same information in natural language, a phenomenon termed "structured but silent."
By Kyojun Choo, Minsoo Song, Yunju Kang, Chanjun Park
The paper introduces EmoNet‑Face‑HQ, a fine‑grained emotion recognition benchmark that uses generated portraits and a 40‑category taxonomy to evaluate vision‑language models (VLMs). It finds that VLMs perform poorly when asked to generate responses but can match or surpass a fine‑tuned model (Empathic‑Insight‑Face) when their logits are read directly as binary queries. The study shows that the benchmark’s difficulty lies in the readout process rather than in perception, and that graded probability outputs yield better performance than simple yes/no questions.
By Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth Andr\'e
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
The paper proposes a production‑centered framework for comparing detection and watermarking techniques across two main AI image and video creation routes: direct visual generation and LLM‑driven code rendering. It introduces an explicit verification specification that separates passive inference, message recovery, and authenticated provenance, and organizes watermarks by production stage for images, videos, source code, and rendering‑aware outputs. The authors outline ten research questions covering identifiability, observability, fair comparison, payload recoverability, reconstruction, synchronization, composition, hybrid local contribution, and private production‑event authentication, and connect the framework to concrete systems such as Claude, OpenAI, and rendering tools.
By Zheng Gao, Xiaoyu Li, Zhicheng Bao, Yang Song, Jiaojiao Jiang