Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits
arXiv:2607. 22564v1 Announce Type: new Abstract: Convolutional neural networks often contain redundant feature maps that increase storage and inference cost.
The paper introduces a structured post‑training pruning method for vision and language transformers called Damage‑Aware Bandit Pruning. It treats the selection of functional units (attention heads and MLP channel groups) as a multi‑armed bandit problem, using paired damage (masked loss minus base loss) as a reward to guide either UCB or Thompson Sampling policies. Experiments on a range of models (GPT‑2, OPT, Pythia, Qwen2.5, SmolLM2, ViT‑B/16, DeiT‑Tiny, Swin‑Tiny) show that the bandit approaches generally reduce degradation compared to budgeted‑greedy baselines, with statistically significant improvements in most comparisons.
arXiv:2607. 22564v1 Announce Type: new Abstract: Convolutional neural networks often contain redundant feature maps that increase storage and inference cost.
arXiv:2606. 07615v1 Announce Type: cross Abstract: Deep neural networks often contain redundant hidden units.
QCPruner is a training‑free visual token pruning method that conditions both token selection and representation on the query via bilateral utility weighting. It fuses keyword‑matched query anchors with cross‑modal cues to compute a nonnegative facility‑location objective that is monotone and submodular, guaranteeing a (1‑1/e) greedy approximation. Across multiple multimodal large language models, QCPruner consistently outperforms existing pruning methods, achieving over 96% of unpruned performance even with very few tokens retained.
arXiv:2607. 09786v1 Announce Type: new Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer.
arXiv:2607. 12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent.
The paper introduces Odds‑Ratio Thompson Sampling (OR‑TS), a method for batched multi‑armed bandits that updates the joint posterior over log‑odds contrasts and refits the shared level in each batch, rather than carrying over absolute reward rates. It presents a Bayesian bandit agent with controls for decay of past evidence and aggressiveness of allocation, and evaluates OR‑TS against traditional absolute‑rate memory across 86 public A/B series and synthetic environments. Results show that when the shared level varies significantly, OR‑TS outperforms absolute‑rate memory, reducing regret and ensuring the best arm receives more traffic, while also handling cases where contrasts themselves shift.
ReNFT is a method that repairs mode collapse in diffusion generators after reward post‑training by internally recalibrating probability mass. It identifies suppressed alternatives through anti‑hub prompts and uses two policy‑dominated routes to generate counterfactual proposals, then applies reward‑based ranking and a joint‑and‑paired NFT update to restore diversity while preserving reward. Experiments on PickScore and GenEval show that ReNFT retains almost all of the original reward while significantly boosting diversity metrics.
arXiv:2608.29765v1 Announce Type: new Abstract: Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average...
arXiv:2607. 25091v1 Announce Type: new Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated.
The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.
The paper introduces patterning, a reweighting technique that adjusts preference pairs based on their susceptibility to bias, to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. Using this method, the authors achieve a +14.2 ± 1.2 percentage point improvement on the RM‑Bench Hard split while maintaining overall accuracy, matching the best reported Hard‑split gain from a comparable model. The study also demonstrates that the learned weights are interpretable, transferable across Gemma variants, and partially effective on Llama 3.1 8B.
arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal...