arXiv Machine Learning

An End-to-End Hybrid Quantum--Classical Sampling Workflow for Discrete Markov Random Fields: A Reproducible Case Study

arXiv:2607. 09893v1 Announce Type: cross Abstract: Sampling from discrete Markov random fields (MRFs) is a hard problem.

arXiv AI
Jun 10

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling

arXiv:2606. 09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution $p^\alpha$ elicits RL-level reasoning from base language models without any parameter updates, but the standard Metropolis--Hastings (MH), a Markov Chain Monte Carlo (MCMC) sampler, is both expensive and slow-mixing.

By Hong Guo, Nianhui Guo, Christoph Meinel, Haojin Yang
arXiv AI
Sep 17

The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

The paper presents a Pareto atlas of LLM inference optimizations, mapping cost, quality, and latency trade‑offs for Qwen2.5‑7B‑Instruct on L4, A100, and H100 GPUs. Using 54 measured configurations and a calibrated simulator, it identifies 18 of 36 setups on the Pareto frontier, showing that combined methods outperform single ones. Quality tests reveal that AWQ 4bit and FP8 weights offer significant latency reductions while largely preserving accuracy, but naive FP8 KV caching fails to answer any questions correctly.

By Srikanta Datta Tumkur, Jay Iyer, Mehar Simhadri, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly
Hugging Face Trending Papers
Jul 5

Asymptotic-Preserving A Posteriori Analysis of Diffusion and Flow-Matching Samplers

Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.

arXiv Machine Learning
Sep 3

Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

The paper introduces a multi‑shell decoder for Leech‑lattice vector quantization, achieving the best reported 2‑bit quality under its evaluation protocol. It presents a GPU‑friendly layout that fuses dequantization with matrix‑vector multiplication, demonstrating significant speed and memory advantages over traditional one‑hot masks and other 4‑bit methods. Experiments show the new kernel outperforms baseline approaches across multiple model sizes, with measurable gains in throughput and reduced byte traffic.

By Pier-Jean Malandrino (Scub)