Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

Hugging Face Trending Papers
3d ago

StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs

StagQ is a multi‑precision weight format for large language models that uses a 2‑bit group‑wise affine base followed by optional 1‑bit refinement planes. Each supported precision can be read as a prefix of the main stream, decoded via a shared affine map without per‑weight lookups, and a sparse side record stores the few weights that the grid handles poorly. Experiments show that StagQ outperforms baseline multi‑precision schemes on Llama‑3.1‑8B, Phi‑4, and OLMo‑2‑7B across various bit‑widths, and its GPU kernel is faster than baseline kernels for most shape‑precision combinations.

Hugging Face Trending Papers
3d ago

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

ThunderSyncRL is a training framework that eliminates idle time in agentic reinforcement learning by starting gradient computation immediately once all necessary inputs are available, thereby avoiding policy staleness. It applies to group relative policy optimization (GRPO) by computing trajectory score gradients as soon as rewards arrive, and to on‑policy distillation (OPD) by updating gradients for completed agentic turns while tool calls execute. Experiments on SWE‑bench Verified and Terminal Bench 4.0 show that ThunderSyncRL matches synchronous training performance up to 1.9× faster and outperforms asynchronous training by up to 2.47 percentage points at a fixed budget.

Hugging Face Trending Papers
3d ago

The Optimization Landscape of Learning Compacted Context Models

The paper investigates the challenges of optimizing compacted context models, particularly the KV cache, in continual learning scenarios. It identifies the optimization landscape as brittle and flat, and proposes a simplified Perceiver-based architecture that matches or surpasses full Perceiver transformers in continuous context compaction. Experiments on MCQ tasks in Finance, Legal, Gutenberg, and Code demonstrate the effectiveness of this approach.

Hugging Face Trending Papers
3d ago

Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation

The paper introduces a game-theoretic approach to text revision, treating token positions as players and vocabulary items as actions, with utilities based on a language model’s log conditional probability. It shows that Nash equilibria can yield exponentially higher likelihoods than autoregressive outputs as sequence length increases, and proposes Nash decoding, an algorithm that finds an ε-Nash equilibrium in O(1/ε) time. Experiments on CLAPNQ, PubMedQA, and CoQA demonstrate that equilibria derived from masked language models achieve higher F1 and ROUGE scores than autoregressive models, up to 18× larger, without fine-tuning, though with extra test-time computation.

arXiv Computer Vision
3d ago

Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild

The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.

By Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt, Manuela Gonz\'alez-Gonz\'alez, Gustave Cortal, Alessandro Lameiras Koerich, Marco Pedersoli, Alain Finkel, Simon Bacon, Eric Granger
arXiv Machine Learning
3d ago

Learning Jazz Pianist Style with Cross-Attention Conditioning

The paper demonstrates that a pretrained symbolic music transformer already encodes jazz pianist identity sufficiently for accurate classification across two benchmarks. By adding cross‑attention over learned pianist embeddings, the model can generate music conditioned on a specific artist’s style, and evaluation protocols confirm that the generated continuations are correctly attributed to the intended pianist. Additionally, the classifier is repurposed to identify the most characteristic moments in a performance, revealing the musical gestures that distinguish each pianist’s voice.

By Drew Edwards, Akira Maezawa, Simon Dixon
arXiv AI
3d ago

Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts

The paper introduces a hardware-software co‑design framework that compresses Mixture‑of‑Experts (MoE) model weights into low‑precision, hardware‑native sparse representations, enabling efficient execution on Sparse Tensor Cores (SpTCs). By relaxing discrete support selection through continuous reparameterization, the method jointly optimizes quantized weights and supports a router‑weighted reconstruction objective, achieving up to 4.35 percentage‑point gains in joint sparse‑quantization accuracy while retaining 96.09% of the original model’s performance. A custom grouped sparse GEMM kernel further boosts inference speed, outperforming NVIDIA’s baseline by up to 1.65× and reducing latency by up to 4.03× on B200 GPUs.

By Kwanhee Lee, Namhoon Lee, Dan Alistarh
arXiv AI
3d ago

Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models

FD‑SCoPE is a language‑model framework that answers clinicians’ questions about systematic review evidence tables, exposing the underlying query, selected trials, and derivation rule for each answer. It handles both directly recorded attributes and derived attributes, achieving high accuracy on an oncology evidence table of 159 immune‑checkpoint inhibitor trials. After incorporating expert corrections, its performance on unseen questions improved from 77.9% to 84.9% F1.

By Manan Roy Choudhury, Suparno Roy Chowdhury, Swastik Sahoo, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani, Mohamad Bassam Sonbol, Irbaz Bin Riaz, Vivek Gupta
arXiv Computer Vision
3d ago

An Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories

The paper introduces the Elastic Shape Variational Autoencoder (ES‑VAE), a geometry‑aware generative model for skeletal pose trajectories that uses the transported square‑root velocity field representation on Kendall's shape manifold to remove rigid transformations and temporal rate variability. ES‑VAE maps sequences to a low‑dimensional latent space via the Riemannian logarithm map and reconstructs them using the exponential map. Experiments on gait analysis for clinical mobility scoring and action recognition on the NTU RGB+D dataset show that ES‑VAE outperforms standard VAEs and several sequence‑modeling baselines.

By Arafat Rahman, Shashwat Kumar, Laura E. Barnes, Anuj Srivastava
arXiv AI
3d ago

IntentCoding: Amplifying User Intent in Code Generation

IntentCoding is a decoding strategy that amplifies user intent in large language model code generation by masking the intent and applying a multi‑strength ensemble mechanism. It is model‑agnostic, requires no extra training, and integrates with existing decoding procedures. Experiments on the new CodeConstraints benchmark and other datasets show significant improvements in constraint satisfaction and functional correctness, with up to 71.0% relative gains on CodeConstraints and 29.3% on HumanEval and LiveCodeBench compared to greedy decoding.

By Zheng Fang, Yihong Dong, Lili Mou, Dongming Jin, Zhi Jin, Ge Li
arXiv Machine Learning
3d ago

Fixed Universal Transformers

The paper introduces fixed universal transformers, which are transformers with immutable internal parameters that can emulate any transformer within a specified class by encoding the target model’s description into the input embedding. The authors provide explicit sparse constructions that achieve universality when the embedding dimension is large enough, and demonstrate that universality is generic—randomly initialized transformers are almost surely universal. Empirical tests on parenthesis balancing and multi‑hop reasoning tasks support the theory, suggesting that a transformer’s expressive power largely stems from its input representation rather than its learned weights.

By Jingwen Liu, Alexandr Andoni, Daniel Hsu
arXiv AI
3d ago

Misinformation Without Triggers: From Factual Answers to Downstream Decisions

The study investigates how false content in training data can influence language models’ downstream decisions even without explicit triggers. By comparing models trained on misleading versus truthful documents in a controlled decision task and a real‑world bushfire case, the authors find a gap between factual answers and the decisions derived from them: correct facts do not always lead to correct decisions, and removing an injected number does not eliminate the misleading narrative. The research highlights that data poisoning can subtly alter model behavior beyond what is detectable through direct probing.

By Lin Tian, Marian-Andrei Rizoiu
arXiv Machine Learning
3d ago

Architecture-Dependent Fusion Pathways in MLLMs

The paper investigates how visual and textual information are fused in Multimodal Large Language Models (MLLMs). By analyzing concatenation and native multimodal architectures through alignment decoupling, attention routing, entropy, intrinsic dimensionality, and causal interventions, the authors uncover two distinct fusion pathways: concatenation models use a text‑first, vision‑later strategy, while native models integrate vision and text earlier and reorganize feature spaces. The study also employs visual CKA to test the Platonic Representation Hypothesis, offering a mechanistic view of multimodal fusion and informing architecture‑aware diagnostics.

By Hebao Zhu, Dongxia Wu
arXiv Computation and Language
3d ago

Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation

The study explores how the language used for reasoning affects retrieval‑augmented generation (RAG) in a monolingual German setting. Using a German RAG question‑answering testbed based on the tabletop game The Dark Eye, the authors show that aligning the reasoning language with the query and retrieved documents improves performance, with German reasoning outperforming French reasoning. However, German reasoning still does not surpass the model’s native English reasoning, indicating that native multilingual reasoning is necessary for optimal results.

By Oliver Hauck, Mario Sanz-Guerrero, Katharina von der Wense
arXiv Machine Learning
3d ago

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

AdaStep introduces an adaptive step-credit weighting technique for agentic reinforcement learning, addressing the coarse granularity of trajectory-level objectives in long-horizon LLM agents. By formulating the weighting as a mean-squared-error estimation problem and deriving an optimal per-state shrinkage coefficient, AdaStep selectively preserves local credit when return variation is due to the chosen action and suppresses it when downstream randomness dominates. The method requires only lightweight scalar computations, no critic or extra rollouts, and demonstrates consistent performance gains across three model backbones on ALFWorld, WebShop, and ScienceWorld.

By Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan
arXiv AI
3d ago

CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation

CONTRA is a training‑free method that discovers and qualifies behavior‑changing questions for selective clarification in large language model (LLM) code generation. It first generates candidate questions, filters out those unrelated to required behavior or already resolved, then creates programs conditioned on two plausible answers to check for stable behavioral differences on shared inputs. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 across four coding agents, outperforming baselines by 13.88 percentage points, and it is also implemented as a Claude Code plugin for practical use.

By Zheng Fang, Yongmin Li, Yichang Zhang, Dongming Jin, Haoyu Wang, Shuai Wang, Zhi Jin, Ge Li
arXiv AI
3d ago

On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

The paper introduces CourseChat, an on‑premises, multi‑course retrieval‑augmented generation (RAG) tutor designed for undergraduate business education. It runs behind a campus web gateway, with each of six courses identified by a unique course reference number (CRN) sharing dual AI hosts that provide a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. The authors evaluate different model sizes, noting that larger models failed speed requirements while a 12B and 7B model met the classroom speed gate; a mixture‑of‑experts variant improved some corrections but introduced new errors, leading them to retain an 8B model for production pending further improvement. The study highlights that model selection, evidence sourcing, serving compatibility, and product design must be considered together, though it does not demonstrate learning gains and calls for separate evaluation of faculty ratings, peak‑load capacity, and public‑gateway acceptance.

By Sidney Shapiro, Joshua Lindemann
arXiv Machine Learning
3d ago

Divergence controls entropy in distillation

The paper investigates how the choice of divergence in knowledge distillation affects the entropy of the student model. It shows that forward KL increases student entropy beyond the teacher’s, while reverse KL decreases it, and that interpolating between the two yields a smooth entropy change early in training but a sharp shift at convergence. The study also finds that on‑policy distillation’s lower entropy stems from token‑level reverse KL rather than sampling, positioning divergence as an implicit entropy regularizer, especially evident in self‑distillation scenarios.

By Nicolas Zucchet, Scott W. Linderman
arXiv Machine Learning
3d ago

GTDD: Generative Test-Driven Development for AI Coding Agents with Adversarial Testing

The paper introduces Generative Test-Driven Development (GTDD), a method where a separate testing agent continuously generates new inputs after each coding agent’s implementation, guided by a human-specified behavioral contract and prior feedback. A trusted evaluator validates these inputs, provides reduced counterexamples, and stores them for regression testing, ensuring that development repeatedly confronts failures beyond the initial examples. Experimental results on a stateful key-value store show that policies regenerating tests during development achieve lower mean failure rates than a single-generation policy, while providing the tester with the candidate’s source code does not yield additional improvement.

By Masahiro Kato
arXiv Machine Learning
3d ago

HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

HyperThink is a text-to-parameter hypernetwork that replaces long-form reasoning traces in large language models with a single, query-conditioned parameter update. The hypernetwork predicts updates to a small subset of the base model’s parameters, and a vector-quantized decoder limits these updates to reusable patterns, improving robustness and transfer. Trained end-to-end on the base model’s outputs, HyperThink eliminates the need for intermediate reasoning traces at test time, producing concise step-by-step solutions with fewer tokens while maintaining strong reasoning performance and improving the accuracy‑latency trade‑off on mathematical and general reasoning tasks.

By Donggyun Kim, Jack Lu, Chanwoo Kim, Mengye Ren, Seunghoon Hong