arXiv Computer Vision

Multi-Scale Temporal Domain Alignment for Federated Video Domain Adaptation

arXiv Computer Vision
Sep 7

Intrinsic Temporal Adaptation of CLIP for Partially Relevant Video Retrieval

The paper introduces Intrinsic Temporal Adaptation (ITA) for Partially Relevant Video Retrieval (PRVR), a task that seeks untrimmed videos containing moments relevant to a text query. ITA employs a Backbone-Internal Temporal Adaptation that lets the final visual transformer layers attend to neighboring frames, creating temporally aware embeddings while keeping CLIP frozen. Additionally, an Affinity-Weighted Gradient Propagation technique softly aggregates top‑k frames based on text‑frame affinities to better handle the weakly supervised nature of PRVR, leading to state‑of‑the‑art performance and more accurate frame‑level evidence retrieval.

By Hyun Seok Seong, Woojin Jun, SuBeen Lee, Jae-Pil Heo
arXiv AI
2d ago

Video Generation Models: A Survey of Post-Training and Alignment

arXiv:2610.00812v1 Announce Type: cross Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamic...

By Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli
arXiv AI
Sep 24

LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations

arXiv:2609.28086v1 Announce Type: cross Abstract: We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings...

By Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur
Hugging Face Trending Papers
Jul 23

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions.

arXiv Computer Vision
Aug 26

Layer-Aware Video Composition via Split-then-Merge

The paper introduces Split-then-Merge (StM), a new framework for generative video composition that improves control and tackles data scarcity. StM divides a large set of unlabeled videos into dynamic foreground and background layers, then self‑composes them to learn how subjects interact with varied scenes. The method employs a transformation‑aware training pipeline with multi‑layer fusion, augmentation, and an identity‑preservation loss, achieving superior performance over state‑of‑the‑art methods in both quantitative and qualitative evaluations.

By Ozgur Kara, Yujia Chen, Ming-Hsuan Yang, James M. Rehg, Wen-Sheng Chu, Du Tran