Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,303 stories · RSS feed

arXiv Machine Learning
2d ago

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
arXiv AI
2d ago

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.

By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li
arXiv AI
2d ago

On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

The paper introduces CourseChat, an on‑premises, multi‑course retrieval‑augmented generation (RAG) tutor designed for undergraduate business education. It runs behind a campus web gateway, with each of six courses identified by a unique course reference number (CRN) sharing dual AI hosts that provide a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. The authors evaluate different model sizes, noting that larger models failed speed requirements while a 12B and 7B model met the classroom speed gate; a mixture‑of‑experts variant improved some corrections but introduced new errors, leading them to retain an 8B model for production pending further improvement. The study highlights that model selection, evidence sourcing, serving compatibility, and product design must be considered together, though it does not demonstrate learning gains and calls for separate evaluation of faculty ratings, peak‑load capacity, and public‑gateway acceptance.

By Sidney Shapiro, Joshua Lindemann
arXiv Machine Learning
2d ago

Divergence controls entropy in distillation

The paper investigates how the choice of divergence in knowledge distillation affects the entropy of the student model. It shows that forward KL increases student entropy beyond the teacher’s, while reverse KL decreases it, and that interpolating between the two yields a smooth entropy change early in training but a sharp shift at convergence. The study also finds that on‑policy distillation’s lower entropy stems from token‑level reverse KL rather than sampling, positioning divergence as an implicit entropy regularizer, especially evident in self‑distillation scenarios.

By Nicolas Zucchet, Scott W. Linderman
arXiv Machine Learning
2d ago

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

The paper introduces Generative Test-Driven Development (GTDD), a method where a separate testing agent continuously generates new inputs after each coding agent’s implementation, guided by a human-specified behavioral contract and prior feedback. A trusted evaluator validates these inputs, provides reduced counterexamples, and stores them for regression testing, ensuring that development repeatedly confronts failures beyond the initial examples. Experimental results on a stateful key-value store show that policies regenerating tests during development achieve lower mean failure rates than a single-generation policy, while providing the tester with the candidate’s source code does not yield additional improvement.

By Masahiro Kato
arXiv Machine Learning
2d ago

HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

HyperThink is a text-to-parameter hypernetwork that replaces long-form reasoning traces in large language models with a single, query-conditioned parameter update. The hypernetwork predicts updates to a small subset of the base model’s parameters, and a vector-quantized decoder limits these updates to reusable patterns, improving robustness and transfer. Trained end-to-end on the base model’s outputs, HyperThink eliminates the need for intermediate reasoning traces at test time, producing concise step-by-step solutions with fewer tokens while maintaining strong reasoning performance and improving the accuracy‑latency trade‑off on mathematical and general reasoning tasks.

By Donggyun Kim, Jack Lu, Chanwoo Kim, Mengye Ren, Seunghoon Hong
arXiv Computation and Language
2d ago

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.

By Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
arXiv AI
2d ago

Distributed Learning with Selective State Space Models: Architecture-Aware Convergence Analysis

The paper investigates how modern selective state space models (SSMs), like Mamba2, behave in distributed learning settings. It derives architecture‑aware gradient and smoothness bounds for single‑ and multi‑layer selective SSMs, and provides convergence bounds for FedAvg and FedProx that incorporate recurrent stability, input‑dependent discretization, and state‑projection norms. Numerical experiments validate the theoretical bounds and compare nine federated learning algorithms on Mamba2 language modeling across six text domains, demonstrating that SSM‑specific bounds help interpret practical federated learning behavior.

By Adam Piaseczny, Md Kamran Chowdhury Shisher, Shiqiang Wang, Christopher G. Brinton
arXiv AI
2d ago

Large Language Continuous Diffusion Models

The paper introduces Sigma, a large-scale continuous diffusion language model (3B/8B parameters) that uses steerable, low-dimensional ODE/SDE latent trajectories to address the non-smoothness of discrete diffusion models. Sigma is trained blockwise via likelihood optimization, jointly denoises Gaussian-corrupted token embeddings, and learns an optimal embedding geometry, leveraging pre-trained autoregressive weights for faster training. During inference, classifier-free guidance and score temperature are identified as essential for high-fidelity reasoning and coding, and Sigma matches or exceeds discrete models on benchmarks such as GSM8K, Minerva, HumanEval, MBPP, MATH-500, and AIME, while also revealing unique structural benefits like embedding-space steering and graceful degradation for low NFEs.

By Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis, Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov, Ante Juki\'c, Arash Vahdat, Morteza Mardani
arXiv AI
2d ago

LEAP: Learning Efficient Action Proposals For LLM Agents

LEAP: Learning Efficient Action Proposals For LLM Agents proposes a method to speed up large language model agents by training a small 0.6B drafter to predict target actions accurately. The approach uses a latency framework that balances drafting, verification, and execution costs, achieving up to 60% faster end‑to‑end wall clock time without reducing task success. LEAP can be online trained, eliminating the need for prior trace collection and making it practical for real‑world deployment.

By Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang
arXiv AI
2d ago

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

DyRA introduces a dynamic, input‑adaptive approach to improve matrix multiplication in deep neural networks by correcting residual output errors during inference. Unlike prior methods that approximate only the weight matrices, DyRA directly optimizes low‑rank factors of the output, combining efficient structured computation with input‑dependent correction. Experiments across vision, speech, and language models show that DyRA consistently enhances the accuracy‑efficiency trade‑off, achieving a 1.5× GPU speedup for DINOv3 while reducing accuracy loss more than threefold compared to weight‑only baselines.

By Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim
arXiv AI
2d ago

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

The paper introduces AgentSecGraph, a static analysis framework that builds a Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation in large language model (LLM) agents. It enriches operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. The authors also present AgentSecBench, a corpus of 67 real-world LLM-agent repositories, and demonstrate that their analyzer identifies thousands of operation candidates, recovers substantial dependency and guard evidence, and distinguishes guarded behaviors from vulnerabilities with high accuracy.

By Hang Cui
arXiv Computation and Language
2d ago

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

The paper presents a method for improving simultaneous speech translation by adapting a full‑utterance speech language model with prefix supervision derived from its own complete and partial waveform translations, eliminating the need for transcripts or human translations. Experiments on FLEURS and CoVoST2 across three language directions show that prefix training enhances quality–latency trade‑offs, especially when combined with multi‑turn append‑only decoding, and that a confidence threshold effectively controls the inference‑time quality–latency balance. The study also explores the impact of synthesis margin on translation quality and calibration, finding a non‑monotonic relationship with latency.

By Hieu Hoang, Amittai Axelrod
arXiv Machine Learning
2d ago

Online Verification of Language Model Responses Under Cost Constraints

The paper introduces OMVV, an online multi-verifier algorithm that maintains a pool of weak verifiers with varying costs and performance. It adaptively selects a verifier each round using an online score combiner and exponential-weights routing, providing distribution-free guarantees on false-accept and false-reject rates. Experiments on reasoning benchmarks show OMVV achieves higher accuracy at lower verification cost than any single fixed verifier across different budgets.

By Erfan Hajihashemi, Yanning Shen
arXiv Machine Learning
2d ago

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Pivot‑SD is an offline self‑distillation framework for masked diffusion language models that focuses training on high‑impact commitments, called pivots, identified by an information‑gain metric. By supervising only these pivots—using cross‑entropy for successful trajectories and targeted unlikelihood for failed ones—Pivot‑SD improves LLaDA‑8B‑Instruct on math and code benchmarks with just 200 questions and four rollouts each.

By Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan
arXiv Machine Learning
2d ago

Planning to Learn

The paper introduces a new loss function called the horizon loss for training classifiers. It argues that the exact policy gradient used in reinforcement learning is myopic, whereas cross‑entropy is patient, and the horizon loss interpolates between the two by truncating the total error at the remaining learning. Experiments on MNIST and ImageNet with ResNet and ViT models show that horizon loss consistently improves top‑1 accuracy over cross‑entropy, especially when label noise is present.

By Ian Osband
arXiv AI
2d ago

Dynamic LLM Routers are Often Misguided

Dynamic LLM routers aim to reduce inference costs by directing each query to the cheapest capable model. In a study of six commercial routers across 14 settings and eight task categories, none surpassed a simple random router that selects between two well-chosen models at the same cost, with some underperforming by over 10 percentage points. The authors identify four common patterns—difficulty blindness, length reversal, semantic matching, and roster suboptimality—that explain this gap and propose a new evaluation method and a simple two-model router that mitigates these patterns, though its advantage over random routing remains modest.

By Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang
arXiv AI
2d ago

OptiSelect: How does the Optimizer Shape Data Curriculum?

OptiSelect is a framework that incorporates the optimizer’s effect into online data selection for large language model pretraining. The study shows that optimizers like Lion and Muon, which use sign-based or polar-tangential preconditioners, suffer from a discriminability collapse that limits selection gains, whereas diagonal‑adaptive optimizers such as AdamW and Sophia can achieve higher gains. Experiments on 124M and 720M models confirm the theory and demonstrate that OptiSelect remains effective even when data is rephrased.

By Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi
arXiv AI
2d ago

Evaluating the Retrieval Robustness of Large Language Models

The paper evaluates how robust large language models (LLMs) are when using retrieval‑augmented generation (RAG) in practical settings. It investigates whether RAG always outperforms non‑RAG approaches, whether adding more retrieved documents helps, and whether the order of documents matters, using a benchmark of 1,891 samples across five datasets and three task categories. Experiments with 11 LLMs show generally high retrieval robustness, but performance varies by task and prompting strategy, indicating that adopting RAG should be considered on a case‑by‑case basis.

By Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang