arXiv Machine Learning

When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting

The paper introduces Pref‑MH, an exact Metropolis‑Hastings sampler that uses only stochastic binary pairwise comparisons from judges to sample from distributions conditioned on desired semantic properties. By linking the MH density ratio to the preference odds of the Bradley‑Terry model, the authors devise an accept/reject rule that guarantees convergence to the target distribution. Experiments on text, molecular, and image generation with large‑language‑model and vision‑language‑model judges show Pref‑MH as a practical, flexible method for conditional sampling when comparative feedback is readily available.

arXiv AI
3d ago

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

Fenchel Tilt Flow Control (FTFC) is a new method for fine‑tuning pretrained generative models to arbitrary preference functions. It decouples utility optimization from model fitting by first learning reward and density‑ratio weights on pretrained samples, then freezing these weights to adjust a diffusion or flow model in a single importance‑weighted stage. The approach supports general f‑divergence penalties, achieves exact duality for concave utilities, and demonstrates up to 20× efficiency gains while outperforming baselines on image and molecule generation tasks.

By Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov
arXiv AI
Aug 28

Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models

The paper introduces Dual-Seed Comparison (DSC), a protocol that uses two independent LLM-generated seeds to reduce systematic bias in probabilistic sampling. DSC constructs a bit sequence from the character-level ordinal values of the seeds, normalizes it into a pseudo-uniform variate, and maps it to the target distribution via the inverse cumulative distribution function. Empirical results show DSC outperforms existing methods in 96% of evaluated settings and enhances distributional control in tasks like MCQ generation and attribute-constrained text-to-image prompting.

By Zihao Guo, Hongtao Lv, Chaoli Zhang, Laiguo Yin, Lei Liu, Yonghui Xu, Lizhen Cui
arXiv AI
Jun 10

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling

arXiv:2606. 09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution $p^\alpha$ elicits RL-level reasoning from base language models without any parameter updates, but the standard Metropolis--Hastings (MH), a Markov Chain Monte Carlo (MCMC) sampler, is both expensive and slow-mixing.

By Hong Guo, Nianhui Guo, Christoph Meinel, Haojin Yang
arXiv Machine Learning
Jul 7

Non-Asymptotic Error Bounds for SMC with Biased Proposals: Application to Conditional Diffusion Sampling

arXiv:2607. 04780v1 Announce Type: cross Abstract: Sequential Monte Carlo (SMC) methods are a natural tool for post-hoc conditioning of pretrained generative models, but in many applications the mutation kernels used by the particle system are biased approximations of an ideal Feynman--Kac flow.

By Stanislas Strasman (SU, LPSM), Gabriel Victorino Cardoso (LPSM), Sylvain Le Corff (LPSM), Vincent Lemaire (LPSM), Antonio Ocello
arXiv Machine Learning
1d ago

Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo

The paper introduces a method for specificity‑aware diffusion steering that suppresses undesired samples while preserving desired ones. By formulating the problem as a target‑design task, it derives a time‑dependent target distribution based on overlap between positive and negative reference distributions, and samples from it using a variance‑reduced Sequential Monte Carlo (SMC) sampler. Experiments on synthetic, class‑contrastive, text‑to‑image, and peptide‑MHC tasks demonstrate reduced mode shift, improved sampling stability, and better suppression of undesired regions compared to negative‑guidance baselines.

By Luran Wang, Linrui Ma, Hannes St\"ark, Regina Barzilay
arXiv Machine Learning
Sep 3

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

HyperMC is a multi‑fidelity hyperparameter tuning framework for stochastic gradient Markov chain Monte Carlo (SGMCMC) that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. It uses successive‑halving brackets to explore a continuous hyperparameter space while progressively refining promising configurations within a fixed computational budget. Robust HyperMC further introduces global grid initialization and elite‑guided local refinement to reduce sensitivity to random candidate generation and noisy evaluations, and theoretical analysis shows that the successive‑halving component selects a near‑optimal configuration with high probability under suitable conditions.

By Ming Tan, Xiyun Jiao