The paper introduces the Logit Refiner, a lightweight autoregressive module that restores intra‑scale dependencies in Visual Autoregressive Models (VAR) by sequentially sampling tokens conditioned on frozen backbone features. This refiner adds only about 10% more parameters and less than 5% of the base model’s training compute, and can be applied to any pretrained VAR checkpoint without retraining. Experiments on ImageNet 256×256 show that the refiner consistently improves generation quality across backbones ranging from 310 M to 2 B parameters, enabling a 1.1 B‑parameter model to outperform a model twice its size, and the method generalizes to text‑to‑image generation, demonstrating that the mean‑field bottleneck is effectively alleviated.
By Meimingwei Li, Stefan Andreas Baumann, Felix Krause, Bj\"orn Ommer
The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.
By Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
arXiv:2603. 12478v2 Announce Type: replace-cross Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven.
By Rujie Wu, Haozhe Zhao, Hai Ci, Yizhou Wang
arXiv:2607. 18042v1 Announce Type: cross Abstract: End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes.
By Lingfeng Zhang, Zhanguang Zhang, Liheng Ma, Tongtong Cao, Yingxue Zhang
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
PRISM is a training‑free framework that efficiently selects visual instruction data for multimodal large language models by addressing the anisotropy in visual feature distributions, which causes a Global Semantic Drift. By implicitly re‑centering visual semantics, PRISM removes the influence of global background features, cutting data‑selection and model‑tuning time to 30% of conventional pipelines while improving performance across eight multimodal and three language benchmarks, achieving a 101.7% relative gain over baseline models.
By Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
arXiv:2606. 13156v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong singleshot spatial grounding, yet lack any mechanism to observe and correct their own predictions.
By Animesh Tripathy, Aswanth Krishnan
LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.
By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes. We first ask a diagnostic question: if the policy is given an expert-trajectory future image as privileged input at training and testing time, is that additional visual evidence useful for choosing the current action?
The paper introduces HB‑SJD, a batched Speculative Jacobi Decoding backend that accelerates visual on‑policy distillation (OPD) by allowing images to advance independently and processing multiple tokens in parallel without a draft model. HB‑SJD switches between Full and Compact execution as images finish, reducing rollout and overall training time while maintaining generation quality. Experiments with LlamaGen demonstrate significant speedups without altering the teacher, distillation objective, or optimization procedure.
By Bingqi Shan, Zhehao Yu, Kenhong Lin, Baoquan Zhang
arXiv:2609.39021v1 Announce Type: new
Abstract: Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly...
By Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang
LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.
By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner