Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
5d ago

Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models

FD‑SCoPE is a language‑model framework that answers clinicians’ questions about systematic review evidence tables, exposing the underlying query, selected trials, and derivation rule for each answer. It handles both directly recorded attributes and derived attributes, achieving high accuracy on an oncology evidence table of 159 immune‑checkpoint inhibitor trials. After incorporating expert corrections, its performance on unseen questions improved from 77.9% to 84.9% F1.

By Manan Roy Choudhury, Suparno Roy Chowdhury, Swastik Sahoo, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol, Irbaz Bin Riaz, Vivek Gupta
arXiv Computer Vision
5d ago

An Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories

The paper introduces the Elastic Shape Variational Autoencoder (ES‑VAE), a geometry‑aware generative model for skeletal pose trajectories that uses the transported square‑root velocity field representation on Kendall's shape manifold to remove rigid transformations and temporal rate variability. ES‑VAE maps sequences to a low‑dimensional latent space via the Riemannian logarithm map and reconstructs them using the exponential map. Experiments on gait analysis for clinical mobility scoring and action recognition on the NTU RGB+D dataset show that ES‑VAE outperforms standard VAEs and several sequence‑modeling baselines.

By Arafat Rahman, Shashwat Kumar, Laura E. Barnes, Anuj Srivastava
arXiv AI
5d ago

IntentCoding: Amplifying User Intent in Code Generation

IntentCoding is a decoding strategy that amplifies user intent in large language model code generation by masking the intent and applying a multi‑strength ensemble mechanism. It is model‑agnostic, requires no extra training, and integrates with existing decoding procedures. Experiments on the new CodeConstraints benchmark and other datasets show significant improvements in constraint satisfaction and functional correctness, with up to 71.0% relative gains on CodeConstraints and 29.3% on HumanEval and LiveCodeBench compared to greedy decoding.

By Zheng Fang, Yihong Dong, Lili Mou, Dongming Jin, Zhi Jin, Ge Li
arXiv Machine Learning
5d ago

Fixed Universal Transformers

The paper introduces fixed universal transformers, which are transformers with immutable internal parameters that can emulate any transformer within a specified class by encoding the target model’s description into the input embedding. The authors provide explicit sparse constructions that achieve universality when the embedding dimension is large enough, and demonstrate that universality is generic—randomly initialized transformers are almost surely universal. Empirical tests on parenthesis balancing and multi‑hop reasoning tasks support the theory, suggesting that a transformer’s expressive power largely stems from its input representation rather than its learned weights.

By Jingwen Liu, Alexandr Andoni, Daniel Hsu
arXiv AI
5d ago

Misinformation Without Triggers: From Factual Answers to Downstream Decisions

The study investigates how false content in training data can influence language models’ downstream decisions even without explicit triggers. By comparing models trained on misleading versus truthful documents in a controlled decision task and a real‑world bushfire case, the authors find a gap between factual answers and the decisions derived from them: correct facts do not always lead to correct decisions, and removing an injected number does not eliminate the misleading narrative. The research highlights that data poisoning can subtly alter model behavior beyond what is detectable through direct probing.

By Lin Tian, Marian-Andrei Rizoiu
arXiv Machine Learning
5d ago

Architecture-Dependent Fusion Pathways in MLLMs

The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.

By Hebao Zhu, Dongxia Wu
arXiv Computation and Language
5d ago

Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation

The study explores how the language used for reasoning affects retrieval‑augmented generation (RAG) in a monolingual German setting. Using a German RAG question‑answering testbed based on the tabletop game The Dark Eye, the authors show that aligning the reasoning language with the query and retrieved documents improves performance, with German reasoning outperforming French reasoning. However, German reasoning still does not surpass the model’s native English reasoning, indicating that native multilingual reasoning is necessary for optimal results.

By Oliver Hauck, Mario Sanz-Guerrero, Katharina von der Wense
arXiv Machine Learning
5d ago

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
arXiv AI
5d ago

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.

By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li
arXiv AI
5d ago

On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

The paper introduces CourseChat, an on‑premises, multi‑course retrieval‑augmented generation (RAG) tutor designed for undergraduate business education. It runs behind a campus web gateway, with each of six courses identified by a unique course reference number (CRN) sharing dual AI hosts that provide a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. The authors evaluate different model sizes, noting that larger models failed speed requirements while a 12B and 7B model met the classroom speed gate; a mixture‑of‑experts variant improved some corrections but introduced new errors, leading them to retain an 8B model for production pending further improvement. The study highlights that model selection, evidence sourcing, serving compatibility, and product design must be considered together, though it does not demonstrate learning gains and calls for separate evaluation of faculty ratings, peak‑load capacity, and public‑gateway acceptance.

By Sidney Shapiro, Joshua Lindemann
arXiv Machine Learning
5d ago

Divergence controls entropy in distillation

The paper investigates how the choice of divergence in knowledge distillation affects the entropy of the student model. It shows that forward KL increases student entropy beyond the teacher’s, while reverse KL decreases it, and that interpolating between the two yields a smooth entropy change early in training but a sharp shift at convergence. The study also finds that on‑policy distillation’s lower entropy stems from token‑level reverse KL rather than sampling, positioning divergence as an implicit entropy regularizer, especially evident in self‑distillation scenarios.

By Nicolas Zucchet, Scott W. Linderman
arXiv Machine Learning
5d ago

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

The paper introduces Generative Test-Driven Development (GTDD), a method where a separate testing agent continuously generates new inputs after each coding agent’s implementation, guided by a human-specified behavioral contract and prior feedback. A trusted evaluator validates these inputs, provides reduced counterexamples, and stores them for regression testing, ensuring that development repeatedly confronts failures beyond the initial examples. Experimental results on a stateful key-value store show that policies regenerating tests during development achieve lower mean failure rates than a single-generation policy, while providing the tester with the candidate’s source code does not yield additional improvement.

By Masahiro Kato
arXiv Machine Learning
5d ago

HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

HyperThink is a text-to-parameter hypernetwork that replaces long-form reasoning traces in large language models with a single, query-conditioned parameter update. The hypernetwork predicts updates to a small subset of the base model’s parameters, and a vector-quantized decoder limits these updates to reusable patterns, improving robustness and transfer. Trained end-to-end on the base model’s outputs, HyperThink eliminates the need for intermediate reasoning traces at test time, producing concise step-by-step solutions with fewer tokens while maintaining strong reasoning performance and improving the accuracy‑latency trade‑off on mathematical and general reasoning tasks.

By Donggyun Kim, Jack Lu, Chanwoo Kim, Mengye Ren, Seunghoon Hong
arXiv Computation and Language
5d ago

CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

CLIMB is a training‑free inference‑time framework for multimodal retrieval‑augmented generation. It builds a compact complementary evidence pool using an MMR‑style objective that balances relevance and redundancy, then refines answers with a confidence‑controlled critic that scores relevance, specificity, and cross‑modal alignment. The method stops refinement when confidence no longer rises, improving performance on Encyclopedic‑VQA and InfoSeek without altering the retriever or language model.

By Hang Gao, Wujiang Xu, Zhixing Zhang, Kai Mei, Jingyi Yang, Dimitris N. Metaxas
arXiv AI
5d ago

Large Language Continuous Diffusion Models

The paper introduces Sigma, a large-scale continuous diffusion language model (3B/8B parameters) that uses steerable, low-dimensional ODE/SDE latent trajectories to address the non-smoothness of discrete diffusion models. Sigma is trained blockwise via likelihood optimization, jointly denoises Gaussian-corrupted token embeddings, and learns an optimal embedding geometry, leveraging pre-trained autoregressive weights for faster training. During inference, classifier-free guidance and score temperature are identified as essential for high-fidelity reasoning and coding, and Sigma matches or exceeds discrete models on benchmarks such as GSM8K, Minerva, HumanEval, MBPP, MATH-500, and AIME, while also revealing unique structural benefits like embedding-space steering and graceful degradation for low NFEs.

By Zhihan Yang, Wei Guo, Jean-Marie Lemercier, Simon Welker, Yonggan Fu, Mohammad Mahdi Kamani, Sajad Norouzi, Julius Berner, Tomas Geffner, Karsten Kreis, Yongxin Chen, Molei Tao, John Thickstun, Pavlo Molchanov, Ante Juki\'c, Arash Vahdat, Morteza Mardani
arXiv AI
5d ago

LEAP: Learning Efficient Action Proposals For LLM Agents

LEAP: Learning Efficient Action Proposals For LLM Agents proposes a method to speed up large language model agents by training a small 0.6B drafter to predict target actions accurately. The approach uses a latency framework that balances drafting, verification, and execution costs, achieving up to 60% faster end‑to‑end wall clock time without reducing task success. LEAP can be online trained, eliminating the need for prior trace collection and making it practical for real‑world deployment.

By Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu, Ce Zhang
arXiv AI
5d ago

DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs

DyRA introduces a dynamic, input‑adaptive approach to improve matrix multiplication in deep neural networks by correcting residual output errors during inference. Unlike prior methods that approximate only the weight matrices, DyRA directly optimizes low‑rank factors of the output, combining efficient structured computation with input‑dependent correction. Experiments across vision, speech, and language models show that DyRA consistently enhances the accuracy‑efficiency trade‑off, achieving a 1.5× GPU speedup for DINOv3 while reducing accuracy loss more than threefold compared to weight‑only baselines.

By Daewon Chae, Hyunwon Chung, Changwoo Lee, Hun-Seok Kim
arXiv AI
5d ago

Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents

The paper introduces AgentSecGraph, a static analysis framework that builds a Security-Aware Agent Dependency Graph (Security-ADG) for each security-sensitive operation in large language model (LLM) agents. It enriches operation identity with agent relevance, source and dependency evidence, trust-boundary context, guard evidence, and external-effect semantics. The authors also present AgentSecBench, a corpus of 67 real-world LLM-agent repositories, and demonstrate that their analyzer identifies thousands of operation candidates, recovers substantial dependency and guard evidence, and distinguishes guarded behaviors from vulnerabilities with high accuracy.

By Hang Cui
arXiv Computation and Language
5d ago

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

The paper presents a method for improving simultaneous speech translation by adapting a full‑utterance speech language model with prefix supervision derived from its own complete and partial waveform translations, eliminating the need for transcripts or human translations. Experiments on FLEURS and CoVoST2 across three language directions show that prefix training enhances quality–latency trade‑offs, especially when combined with multi‑turn append‑only decoding, and that a confidence threshold effectively controls the inference‑time quality–latency balance. The study also explores the impact of synthesis margin on translation quality and calibration, finding a non‑monotonic relationship with latency.

By Hieu Hoang, Amittai Axelrod
arXiv Machine Learning
5d ago

Online Verification of Language Model Responses Under Cost Constraints

The paper introduces OMVV, an online multi-verifier algorithm that maintains a pool of weak verifiers with varying costs and performance. It adaptively selects a verifier each round using an online score combiner and exponential-weights routing, providing distribution-free guarantees on false-accept and false-reject rates. Experiments on reasoning benchmarks show OMVV achieves higher accuracy at lower verification cost than any single fixed verifier across different budgets.

By Erfan Hajihashemi, Yanning Shen