Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv Computer Vision
1d ago

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.

By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
arXiv Machine Learning
1d ago

Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers

The paper studies reinforcement learning (RL) for post‑training language models on reasoning tasks, focusing on a hierarchical reward structure where each component becomes relevant only after the previous ones are resolved. It demonstrates that a Transformer‑based actor–critic algorithm that alternates between KL‑regularized policy sampling, critic fitting, and policy updates achieves minimax‑optimal rates in query budget and regularization strength, and is optimal for a fixed number of prompts. In contrast, sampling from a fixed reference distribution, as used in offline reward modeling, only yields a logarithmic regret decay, highlighting the advantage of on‑policy exploration in concentrating on high‑reward regions.

By Naoki Nishikawa, Taiji Suzuki
arXiv Machine Learning
1d ago

Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings

The paper introduces a Hybrid Cross-Modal Attention Network (HCMAN) that fuses mammogram images with structured clinical data using transformer-based cross‑modal attention. Trained on a locally collected dataset of 2,560 images from 1,024 Ethiopian patients, the model achieves 97.8% accuracy, 97.2% sensitivity, 98.3% specificity, and an AUC of 0.987, outperforming image‑only baselines and maintaining robustness to low‑quality images. Its lightweight design allows inference in under two seconds on a standard CPU, making it suitable for deployment in resource‑limited clinical settings.

By Simon Hadush Nrea (Mekelle University, Mekelle, Ethiopia), Filimon Gidey Gebremichael (Mekelle University, Mekelle, Ethiopia), Gebrekirstos Hagos Gebrekirstos (Clinical Oncologist London School of Hygiene and Tropical Medicine London, UK), Yaecob Girmay Gezahegn (Mekelle University, Mekelle, Ethiopia)
arXiv AI
1d ago

Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.

By Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong
arXiv AI
1d ago

Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction

The paper introduces SAGA, a framework that enables large language model agents to evolve by abstracting experiences into reusable principles, procedures, and episodic descriptions. SAGA transforms interaction trajectories into hierarchical knowledge with explicit applicability conditions, linking them back to execution evidence. Experiments on ScienceWorld and ALFWorld show that this execution–abstraction feedback loop improves task performance, and ablation studies confirm the importance of contextual instantiation and action regulation.

By Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li
arXiv AI
1d ago

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.

By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv AI
1d ago

Understanding and Mitigating Inference-Time Overreliance Using Agentic Memory

The paper investigates how large language model agents can over-rely on agentic memory, a phenomenon where retrieved memories distort inference even when they are correctly stored and retrieved. It shows that memory is helpful when past experience fully transfers to the current task but becomes misleading under partial query-memory overlap, a pattern confirmed by controlled experiments. To address this, the authors propose MEMTRIM, a plug‑and‑play framework that indexes memory evidence at write time and limits its reuse at read time, thereby reducing over-reliance without retraining and preserving useful memory benefits across models and memory architectures.

By Luoxi Tang, Yuqiao Meng, Nilesh Auradkar, Muchao Ye, Dazheng Zhang, Zhaohan Xi
arXiv AI
1d ago

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

The paper investigates whether semantic entropy—a measure of disagreement among a language model’s sampled answers—can serve as a cheap signal for deciding when to route a query from a small to a larger language model. Experiments on GSM8K and other benchmarks show that semantic entropy can distinguish small‑model mistakes and improve routed accuracy, but the authors also reveal that a simple question‑difficulty rule can mimic its performance and that other factors (definition of success, benchmark design, live sampling cost) can undermine its effectiveness. They propose a checklist of checks to validate escalation signals and demonstrate how to predict when a cached‑outcome approach will fail. "whyItMatters":"The study highlights the importance of rigorous evaluation of escalation signals, showing that seemingly promising metrics can be misleading without proper controls and that practical routing decisions must account for cost and benchmark design."

By Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan
arXiv AI
1d ago

PsyCIDRA: A Dual-Agent Framework for Psychiatric Interviewing and Diagnostic Reasoning

PsyCIDRA is a dual‑agent framework that couples a free‑form psychiatric interviewer with a diagnostic reasoning agent to support expert review. The interviewer agent uses tools to keep working notes, load expert skills, and pull ICD‑11 references, while the diagnostic agent receives the interview transcript and generates hypotheses with supporting, conflicting, and missing evidence, withholding a final hypothesis if insufficient support exists. In simulations and a blinded human study, PsyCIDRA achieved higher diagnostic agreement and rank‑1 accuracy than direct prompting, indicating its promise for assisting psychiatric assessment through interactive dialogue.

By Milad Mohammadi, Fatemeh Akrami Shamsabadi, Zahra Mohseni, Amirhossein Safdarian, Malekfarhad Malek, Hadi Moradi, Hesham Faili
arXiv Machine Learning
1d ago

Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models

The study investigates whether particular attention heads and individual neurons within those heads in language models are responsible for detecting network infrastructure information—specifically hostnames paired with IP addresses. Using causal ablation and selective testing across five models from three architecture families, the authors find that a small subset of heads reliably identifies such information with near-perfect accuracy. However, the extent to which this responsibility is concentrated in a single neuron varies by model; in some cases a single neuron suffices, while in others the signal is distributed across the head. The findings generalize to an independent reverse‑DNS dataset, though single‑neuron detectors are less robust.

By Abdul Kadir (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence), Md Mohasin Hossain (German Research Center for Artificial Intelligence, Saarland University, Saarbrucken, Germany), Daniel Sonntag (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence)