arXiv AI

Damage-Aware Bandit Pruning for Vision and Language Transformers

The paper introduces a structured post‑training pruning method for vision and language transformers called Damage‑Aware Bandit Pruning. It treats the selection of functional units (attention heads and MLP channel groups) as a multi‑armed bandit problem, using paired damage (masked loss minus base loss) as a reward to guide either UCB or Thompson Sampling policies. Experiments on a range of models (GPT‑2, OPT, Pythia, Qwen2.5, SmolLM2, ViT‑B/16, DeiT‑Tiny, Swin‑Tiny) show that the bandit approaches generally reduce degradation compared to budgeted‑greedy baselines, with statistically significant improvements in most comparisons.

arXiv Computer Vision
Sep 18

QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

QCPruner is a training‑free visual token pruning method that conditions both token selection and representation on the query via bilateral utility weighting. It fuses keyword‑matched query anchors with cross‑modal cues to compute a nonnegative facility‑location objective that is monotone and submodular, guaranteeing a (1‑1/e) greedy approximation. Across multiple multimodal large language models, QCPruner consistently outperforms existing pruning methods, achieving over 96% of unpruned performance even with very few tokens retained.

By Shengli He (Guizhou University), Yongchao Liang (Guizhou University), Roumeng He (Shanghai Ocean University), Junjie Zeng (Guizhou University), Jiyuan He (Guizhou University), Can Wu (Guizhou University), Li Zheng (Guizhou University)
arXiv AI
Jul 15

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

arXiv:2607. 12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent.

By Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang
arXiv Machine Learning
Sep 18

Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits

The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.

By Sulgi Kim
arXiv AI
Sep 2

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

ReNFT is a method that repairs mode collapse in diffusion generators after reward post‑training by internally recalibrating probability mass. It identifies suppressed alternatives through anti‑hub prompts and uses two policy‑dominated routes to generate counterfactual proposals, then applies reward‑based ranking and a joint‑and‑paired NFT update to restore diversity while preserving reward. Experiments on PickScore and GenEval show that ReNFT retains almost all of the original reward while significantly boosting diversity metrics.

By Yuchen Bao, Chao Wen, Haowei Wang, Ruoxin Chen, Donghao Luo, Jiahui Zhan, Wenjian Huang, Shen Chen, Yiting Wang, Taiping Yao, Chengjie Wang, Shouhong Ding, Jianguo Zhang
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Machine Learning
Sep 2

Patterning in Practice: Debiasing Reward Models with Susceptibilities

The paper introduces patterning, a reweighting technique that adjusts preference pairs based on their susceptibility to bias, to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. Using this method, the authors achieve a +14.2 ± 1.2 percentage point improvement on the RM‑Bench Hard split while maintaining overall accuracy, matching the best reported Hard‑split gain from a comparable model. The study also demonstrates that the learned weights are interpretable, transferable across Gemma variants, and partially effective on Llama 3.1 8B.

By George Wang, Elizabeth Donoway, Daniel Murfet