arXiv Computer Vision

Separators in Enhancing Autoregressive Pretraining for Vision Mamba

arXiv Computer Vision
Aug 27

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

The paper introduces MTAR, a training framework for autoregressive image generation that enhances performance through multi-token prediction, token-level contrastive regularization, and semantic dropping. These components address sparse supervision, improve representation discriminability, and accelerate training without affecting inference. On ImageNet, MTAR outperforms LlamaGen with lower FID and faster training, achieving comparable results in only a third of the iterations.

By Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu
arXiv AI
Aug 28

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.

By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny
arXiv Computer Vision
Sep 17

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

FlashAR is a lightweight post‑training adaptation framework that converts a pre‑trained raster‑scan autoregressive image model into a highly parallel generator using two‑way next‑token prediction. It preserves the original training objective by keeping the horizontal head for row‑wise prediction and adding a lightweight vertical head for column‑wise prediction, with a learnable fusion gate to combine the two predictions. A two‑stage adaptation pipeline—first initializing the vertical head from the pre‑trained model and then jointly fine‑tuning—yields up to a 22.9× speedup for 512×512 image generation while using only 0.05% of the original training data.

By Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang
arXiv AI
Aug 28

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

PACE introduces a training‑free Condense‑and‑Extract framework that speeds up Vision‑Language Model inference by first adaptively downsampling visual inputs before encoding and then selectively retaining essential tokens during decoding. The Adaptive Pixel Compressor (APC) reduces encoder workload while preserving global context, and the Dynamic Dual‑Attention Extractor (DDAE) keeps task‑critical details by fusing visual and language signals. Applied to Qwen2.5‑VL‑7B, PACE maintains 93.8% of performance using only 10% of visual tokens, achieving a 3.1× speedup in time to first token.

By Junjie Liu, Shengyuan Ye, Xu Chen