arXiv AI By Yueyang Cang, Xiaoteng Zhang, Zhiyuan Ning, Yuchen He, Li Shi

FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking

Read the original on arXiv AI →

arXiv:2608. 09369v1 Announce Type: cross Abstract: Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 11

LoopVAE: Recurrent Depth Across Scales for Visual Tokenization

LoopVAE introduces a recurrent depth architecture that reuses a scale‑ and loop‑conditioned core across different spatial scales while keeping resolution‑changing transitions separate. The four‑block core applies 28 block operations per encoder or decoder, enabling a 29M‑parameter convolutional model to achieve 0.28 rFID and 32.54 dB PSNR on ImageNet‑256 with roughly 65% fewer parameters than comparable VAEs. Experiments with both convolutional and Transformer operators, as well as ablations on parameter sharing, demonstrate competitive image quality metrics and reveal how targeted loop interventions and truncation affect reconstruction quality and computational trade‑offs.

By Zhiying Lu
arXiv Computer Vision
Sep 21

MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

MT‑WAM enhances the Fast‑WAM framework by adding complementary supervision for future 2‑D point trajectories and visual features while keeping the original training objectives. A lightweight dual‑stream branch and structured attention mask isolate motion‑specific processing, and motion‑stream tokens provide additional dynamics cues to the action expert. During inference, MT‑WAM skips future‑video prediction, using cached video and motion information to achieve higher success rates on LIBERO, LIBERO‑Plus, RoboTwin 2.0 Clean2Rand, and several real‑world tasks.

By Yiguang Yang, Jiankun Peng, Xiaoming Wang, Yiran Zhang, Zhibo Fang
arXiv Computer Vision
Aug 27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU‑TTT is a causal test‑time training method for long‑video multimodal large language models that inserts a convolutional resampler with fast‑weight updates between the vision encoder and the LLM. The fast weights adapt per video and contextualize frame features before compression, while a hybrid selector keeps explicit visual evidence for downstream reasoning. Experiments show that TTT‑Conv outperforms TTT‑MLP and bidirectional Mamba2 on MLVU, and beats attention‑ and fixed‑state recurrent resamplers on three benchmarks, achieving competitive results on five video‑understanding tasks after reducing 512 frames to 128 LLM frames.

By Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny