The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.
By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
The paper studies reinforcement learning (RL) for post‑training language models on reasoning tasks, focusing on a hierarchical reward structure where each component becomes relevant only after the previous ones are resolved. It demonstrates that a Transformer‑based actor–critic algorithm that alternates between KL‑regularized policy sampling, critic fitting, and policy updates achieves minimax‑optimal rates in query budget and regularization strength, and is optimal for a fixed number of prompts. In contrast, sampling from a fixed reference distribution, as used in offline reward modeling, only yields a logarithmic regret decay, highlighting the advantage of on‑policy exploration in concentrating on high‑reward regions.
By Naoki Nishikawa, Taiji Suzuki
The paper introduces a Hybrid Cross-Modal Attention Network (HCMAN) that fuses mammogram images with structured clinical data using transformer-based cross‑modal attention. Trained on a locally collected dataset of 2,560 images from 1,024 Ethiopian patients, the model achieves 97.8% accuracy, 97.2% sensitivity, 98.3% specificity, and an AUC of 0.987, outperforming image‑only baselines and maintaining robustness to low‑quality images. Its lightweight design allows inference in under two seconds on a standard CPU, making it suitable for deployment in resource‑limited clinical settings.
By Simon Hadush Nrea (Mekelle University, Mekelle, Ethiopia), Filimon Gidey Gebremichael (Mekelle University, Mekelle, Ethiopia), Gebrekirstos Hagos Gebrekirstos (Clinical Oncologist London School of Hygiene and Tropical Medicine London, UK), Yaecob Girmay Gezahegn (Mekelle University, Mekelle, Ethiopia)
arXiv:2610.08586v1 Announce Type: new
Abstract: Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complet...
By Sujato Dutta, Sreekruthy Tummala, Shashank Vanga, Ayushmi Pavani
The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.
By Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong
The paper introduces SAGA, a framework that enables large language model agents to evolve by abstracting experiences into reusable principles, procedures, and episodic descriptions. SAGA transforms interaction trajectories into hierarchical knowledge with explicit applicability conditions, linking them back to execution evidence. Experiments on ScienceWorld and ALFWorld show that this execution–abstraction feedback loop improves task performance, and ablation studies confirm the importance of contextual instantiation and action regulation.
By Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li
The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.
By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
The paper investigates how large language model agents can over-rely on agentic memory, a phenomenon where retrieved memories distort inference even when they are correctly stored and retrieved. It shows that memory is helpful when past experience fully transfers to the current task but becomes misleading under partial query-memory overlap, a pattern confirmed by controlled experiments. To address this, the authors propose MEMTRIM, a plug‑and‑play framework that indexes memory evidence at write time and limits its reuse at read time, thereby reducing over-reliance without retraining and preserving useful memory benefits across models and memory architectures.
By Luoxi Tang, Yuqiao Meng, Nilesh Auradkar, Muchao Ye, Dazheng Zhang, Zhaohan Xi
The paper investigates whether semantic entropy—a measure of disagreement among a language model’s sampled answers—can serve as a cheap signal for deciding when to route a query from a small to a larger language model. Experiments on GSM8K and other benchmarks show that semantic entropy can distinguish small‑model mistakes and improve routed accuracy, but the authors also reveal that a simple question‑difficulty rule can mimic its performance and that other factors (definition of success, benchmark design, live sampling cost) can undermine its effectiveness. They propose a checklist of checks to validate escalation signals and demonstrate how to predict when a cached‑outcome approach will fail.
"whyItMatters":"The study highlights the importance of rigorous evaluation of escalation signals, showing that seemingly promising metrics can be misleading without proper controls and that practical routing decisions must account for cost and benchmark design."
By Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan
PsyCIDRA is a dual‑agent framework that couples a free‑form psychiatric interviewer with a diagnostic reasoning agent to support expert review. The interviewer agent uses tools to keep working notes, load expert skills, and pull ICD‑11 references, while the diagnostic agent receives the interview transcript and generates hypotheses with supporting, conflicting, and missing evidence, withholding a final hypothesis if insufficient support exists. In simulations and a blinded human study, PsyCIDRA achieved higher diagnostic agreement and rank‑1 accuracy than direct prompting, indicating its promise for assisting psychiatric assessment through interactive dialogue.
By Milad Mohammadi, Fatemeh Akrami Shamsabadi, Zahra Mohseni, Amirhossein Safdarian, Malekfarhad Malek, Hadi Moradi, Hesham Faili
arXiv:2610.07570v1 Announce Type: new
Abstract: In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism...
By Xiaoyang Wang, Tianrui Wang, Christopher C. Yang
arXiv:2610.08216v1 Announce Type: new
Abstract: Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pool...
By Srikar Alla, Ali Shiri Sichani, Chi-Ren Shyu
arXiv:2610.08231v1 Announce Type: new
Abstract: NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preser...
By Neriah Ben David, Ori Meir, Or Ordentlich
arXiv:2610.08246v1 Announce Type: new
Abstract: Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing plannin...
By Andr\'{e} G. Pereira, Augusto B. Corr\^ea, Felipe Meneguzzi, Jendrik Seipp
arXiv:2610.08250v1 Announce Type: new
Abstract: Large language models are increasingly used to simulate clients for counselor training and psychological counseling research, but reliable simulation r...
By Shixin Peng, Kun Jiang, Jiaxing Zheng, Qihao Yang, Jingying Chen
arXiv:2610.08312v1 Announce Type: new
Abstract: Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, rec...
By Maoqi Liu, Quan Fang, Yufei He
arXiv:2610.08319v1 Announce Type: new
Abstract: In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails...
By Hanchao Zhou, Jialei Li
The study investigates whether particular attention heads and individual neurons within those heads in language models are responsible for detecting network infrastructure information—specifically hostnames paired with IP addresses. Using causal ablation and selective testing across five models from three architecture families, the authors find that a small subset of heads reliably identifies such information with near-perfect accuracy. However, the extent to which this responsibility is concentrated in a single neuron varies by model; in some cases a single neuron suffices, while in others the signal is distributed across the head. The findings generalize to an independent reverse‑DNS dataset, though single‑neuron detectors are less robust.
By Abdul Kadir (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence), Md Mohasin Hossain (German Research Center for Artificial Intelligence, Saarland University, Saarbrucken, Germany), Daniel Sonntag (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence)
arXiv:2610.07247v1 Announce Type: new
Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for tr...
By Heng Liang, Xinwen Zhang, Hongchang Gao
arXiv:2610.06889v1 Announce Type: cross
Abstract: We study the application of large language models (LLMs) to the visual exploration of textual corpora. We introduce zero-shot visualization (ZSV), a...
By Arnau Bueno Tricas, Jose A. Rodr\'iguez-Serrano