Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv Computation and Language
3d ago

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.

By Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
arXiv AI
3d ago

Large Language Continuous Diffusion Models

The paper introduces Sigma, a large-scale continuous diffusion language model (3B/8B parameters) that uses steerable, low-dimensional ODE/SDE latent trajectories to address the non-smoothness of discrete diffusion models. Sigma is trained blockwise via likelihood optimization, jointly denoises Gaussian-corrupted token embeddings, and learns an optimal embedding geometry, leveraging pre-trained autoregressive weights for faster training. During inference, classifier-free guidance and score temperature are identified as essential for high-fidelity reasoning and coding, and Sigma matches or exceeds discrete models on benchmarks such as GSM8K, Minerva, HumanEval, MBPP, MATH-500, and AIME, while also revealing unique structural benefits like embedding-space steering and graceful degradation for low NFEs.

By Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis, Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov, Ante Juki\'c, Arash Vahdat, Morteza Mardani
arXiv AI
3d ago

LEAP: Learning Efficient Action Proposals For LLM Agents

LEAP: Learning Efficient Action Proposals For LLM Agents proposes a method to speed up large language model agents by training a small 0.6B drafter to predict target actions accurately. The approach uses a latency framework that balances drafting, verification, and execution costs, achieving up to 60% faster end‑to‑end wall clock time without reducing task success. LEAP can be online trained, eliminating the need for prior trace collection and making it practical for real‑world deployment.

By Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang
arXiv AI
3d ago

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

DyRA introduces a dynamic, input‑adaptive approach to improve matrix multiplication in deep neural networks by correcting residual output errors during inference. Unlike prior methods that approximate only the weight matrices, DyRA directly optimizes low‑rank factors of the output, combining efficient structured computation with input‑dependent correction. Experiments across vision, speech, and language models show that DyRA consistently enhances the accuracy‑efficiency trade‑off, achieving a 1.5× GPU speedup for DINOv3 while reducing accuracy loss more than threefold compared to weight‑only baselines.

By Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim
arXiv AI
3d ago

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

The paper introduces AgentSecGraph, a static analysis framework that builds a Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation in large language model (LLM) agents. It enriches operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. The authors also present AgentSecBench, a corpus of 67 real-world LLM-agent repositories, and demonstrate that their analyzer identifies thousands of operation candidates, recovers substantial dependency and guard evidence, and distinguishes guarded behaviors from vulnerabilities with high accuracy.

By Hang Cui
arXiv Computation and Language
3d ago

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

The paper presents a method for improving simultaneous speech translation by adapting a full‑utterance speech language model with prefix supervision derived from its own complete and partial waveform translations, eliminating the need for transcripts or human translations. Experiments on FLEURS and CoVoST2 across three language directions show that prefix training enhances quality–latency trade‑offs, especially when combined with multi‑turn append‑only decoding, and that a confidence threshold effectively controls the inference‑time quality–latency balance. The study also explores the impact of synthesis margin on translation quality and calibration, finding a non‑monotonic relationship with latency.

By Hieu Hoang, Amittai Axelrod
arXiv Machine Learning
3d ago

Online Verification of Language Model Responses Under Cost Constraints

The paper introduces OMVV, an online multi-verifier algorithm that maintains a pool of weak verifiers with varying costs and performance. It adaptively selects a verifier each round using an online score combiner and exponential-weights routing, providing distribution-free guarantees on false-accept and false-reject rates. Experiments on reasoning benchmarks show OMVV achieves higher accuracy at lower verification cost than any single fixed verifier across different budgets.

By Erfan Hajihashemi, Yanning Shen
arXiv Machine Learning
3d ago

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot‑SD is an offline self‑distillation framework for masked diffusion language models that focuses training on high‑impact commitments, called pivots, identified by an information‑gain metric. By supervising only these pivots—using cross‑entropy for successful trajectories and targeted unlikelihood for failed ones—Pivot‑SD improves LLaDA‑8B‑Instruct on math and code benchmarks with just 200 questions and four rollouts each.

By Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan
arXiv Machine Learning
3d ago

Planning to Learn

The paper introduces a new loss function called the horizon loss for training classifiers. It argues that the exact policy gradient used in reinforcement learning is myopic, whereas cross‑entropy is patient, and the horizon loss interpolates between the two by truncating the total error at the remaining learning. Experiments on MNIST and ImageNet with ResNet and ViT models show that horizon loss consistently improves top‑1 accuracy over cross‑entropy, especially when label noise is present.

By Ian Osband
arXiv AI
3d ago

Dynamic LLM Routers are Often Misguided

Dynamic LLM routers aim to reduce inference costs by directing each query to the cheapest capable model. In a study of six commercial routers across 14 settings and eight task categories, none surpassed a simple random router that selects between two well-chosen models at the same cost, with some underperforming by over 10 percentage points. The authors identify four common patterns—difficulty blindness, length reversal, semantic matching, and roster suboptimality—that explain this gap and propose a new evaluation method and a simple two-model router that mitigates these patterns, though its advantage over random routing remains modest.

By Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang
arXiv AI
3d ago

OptiSelect: How does the Optimizer Shape Data Curriculum?

OptiSelect is a framework that incorporates the optimizer’s effect into online data selection for large language model pretraining. The study shows that optimizers like Lion and Muon, which use sign-based or polar-tangential preconditioners, suffer from a discriminability collapse that limits selection gains, whereas diagonal‑adaptive optimizers such as AdamW and Sophia can achieve higher gains. Experiments on 124M and 720M models confirm the theory and demonstrate that OptiSelect remains effective even when data is rephrased.

By Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi
arXiv AI
3d ago

Evaluating the Retrieval Robustness of Large Language Models

The paper evaluates how robust large language models (LLMs) are when using retrieval‑augmented generation (RAG) in practical settings. It investigates whether RAG always outperforms non‑RAG approaches, whether adding more retrieved documents helps, and whether the order of documents matters, using a benchmark of 1,891 samples across five datasets and three task categories. Experiments with 11 LLMs show generally high retrieval robustness, but performance varies by task and prompting strategy, indicating that adopting RAG should be considered on a case‑by‑case basis.

By Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang
arXiv AI
3d ago

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.

By Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
arXiv Computation and Language
3d ago

OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models

OLMo-Detect is a new benchmark for membership inference attacks (MIAs) on large language models, covering pre‑training, mid‑training, and post‑training stages. It aligns member and non‑member samples on key axes and rigorously filters non‑members using infini‑gram, also offering a shifted variant to test distribution robustness. Evaluation of 15 unsupervised and 3 supervised MIAs on the OLMo 2 family shows limited overall performance (best AUC 0.68), peak detection at mid‑training driven by data type, and sensitivity to distribution shifts, with findings generalizing to OLMo 3 and other models.

By Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau
arXiv Computer Vision
3d ago

Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs

Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.

By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun
arXiv Computation and Language
3d ago

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.

By Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne
arXiv Machine Learning
3d ago

Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild

The paper investigates whether adding a short classification instruction after a user’s prompt improves the ability of activation probes to detect malicious inputs in large language models. Across 13 safety benchmarks and three open‑weight model families, a classification suffix consistently boosts out‑of‑distribution detection (up to ~4 AUC points) compared to no suffix, and the benefit transfers to multi‑position pooling probes used in production. The improvement stems from the classification format itself rather than the specific content of the instruction, though the optimal suffix varies with the model and readout type.

By Elad David, Max Fomin
arXiv AI
3d ago

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

NeutronGym is the first executable environment that lets language‑model agents design neutron instruments, using tools that validate their builds, McStas ray‑tracing for simulation, and a level‑resolved grading ladder that evaluates syntax, runtime, structure, and science without an LLM judge. The platform provides procedural families of instrument layouts with held‑out parameter regimes and a curated set of 16 tasks from published instruments, called McStasBench, which includes memorization probes and a sandbox. Experiments show that while seven models can reproduce at most seven of the 16 tasks and none reaches a reference design, reinforcement learning can dramatically improve performance—e.g., Qwen3‑8B’s success rate jumps from 11% to 77% on held‑out instances—highlighting the ladder’s importance for partial credit and the potential of RL to match classical optimizers under realistic simulation budgets.

By Lijie Ding, Changwoo Do
arXiv Computation and Language
3d ago

RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation

RT‑SFT is a method for text style transfer that uses roundtrip translation through a pivot language to strip stylistic information from monolingual corpora, creating pseudo‑parallel data. This data is then used to LoRA‑finetune an instruction‑tuned large language model as a stylizer, allowing the model to rewrite sentences in a target style while preserving meaning. Experiments across four style domains show that RT‑SFT surpasses state‑of‑the‑art approaches, including few‑shot in‑context learning, and offers effective retrieval augmentation for expert style domains with strict terminology.

By Ruoxi Liu, Philipp Koehn