arXiv AI

Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment

arXiv:2606. 27731v1 Announce Type: cross Abstract: Despite their strong general capabilities, large language models (LLMs) often remain unreliable when outputs must be numerically precise.

arXiv AI
Sep 10

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

The paper argues that relying solely on zero‑shot task accuracy is insufficient for evaluating quantized large language models (LLMs) because accuracy ignores changes in the full predictive distribution. It proposes a distribution‑sensitive framework that measures fidelity loss by computing statistical distances—such as Jensen‑Shannon Divergence and Total Variation Distance—between the full‑vocabulary output distributions of a full‑precision BF16 reference and its quantized counterparts. Experiments across five foundation architectures and four reasoning benchmarks show that these divergence metrics increase with stronger quantization, revealing distributional drift that top‑1 accuracy fails to capture, and suggest that mixed‑precision Q4_K schemes can offer lower divergence than uniform Q4_0 at comparable memory usage.

By Shahzeb Qamar, Lorenz Sparrenberg, Christian Bauckhage, Baha Rababah, Carson Leung, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Rafet Sifa
arXiv AI
Jul 28

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

arXiv:2607. 24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling.

By Xiaoyi Jiang, Jingyuan Li, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu
arXiv Machine Learning
Aug 20

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

The paper introduces the concept of decision‑metric alignment, which ensures that Euclidean distance to a goal latent in JEPA‑style latent world models correctly ranks action sequences for model‑predictive control. It proposes two metrics—Plan‑Real Spearman and CEM‑stage Spearman—to evaluate latent–real rank agreement, and identifies encoder distortion, terminal rollout error, and candidate margins as key factors affecting alignment. Building on these insights, the authors present DA‑LeWM, an enhanced latent world model that incorporates inverse‑dynamics and demonstration‑conditioned goal‑action heads, leading to faster convergence and higher online success rates compared to the baseline LeWM while maintaining similar probe scores.

By Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
arXiv AI
3d ago

Fast Generalized Neural Tangent Kernel Statistics via Trace Estimation

The paper presents efficient approximations for key statistics of the Neural Tangent Kernel (NTK) in finite-width neural networks using randomized trace estimation (Hutch++). It demonstrates that the NTK trace, Frobenius norm, effective rank, and alignment can be estimated with high accuracy via matrix-free products, leveraging the NTK’s positive-semidefinite structure to use one-sided estimators with forward or reverse-mode differentiation. Experiments on MLPs, GRUs, and a 410‑million‑parameter Transformer show orders‑of‑magnitude speedups and enable practical state‑space NTK diagnostics at large scales, including applications to RNN training and data‑scarce knowledge distillation.

By James Hazelden, Balaaji Reddy Nagireddy, Eric Shea-Brown
Hugging Face Trending Papers
Aug 19

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

The paper investigates how latent world models (specifically JEPA-style models) use Euclidean distance to a goal latent as a cost for model‑predictive control (MPC). It introduces two metrics—Plan‑Real Spearman and CEM‑stage Spearman—to evaluate how well latent‑space distances align with real‑task progress, a property termed decision‑metric alignment. By identifying encoder distortion, terminal rollout error, and candidate margins as key factors, the authors propose DA‑LeWM, which augments the base model with inverse‑dynamics and demonstration‑conditioned goal‑action heads, leading to faster convergence and higher online success while maintaining similar probe scores.

arXiv AI
Aug 25

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

The paper introduces the Capability-Driven Multimodal Scaling Law, a cross-family framework that predicts vision-language model (VLM) benchmark accuracy from a low-dimensional textual capability score extracted via PCA. By training over 150 VLMs on 34 large language models across seven families, the authors demonstrate that the law accurately extrapolates transfer rates from 8B to 72B‑parameter backbones, predicts full training trajectories, and generalizes to unseen model families. The study also reveals actionable insights, such as certain textual benchmarks negatively correlating with multimodal performance and base LLMs outperforming instruction-tuned counterparts as VLM backbones due to higher absorption rates.

By Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen, Jingang Wang, Xunliang Cai
arXiv Computation and Language
Sep 2

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani