arXiv AI

DIMOS: Disentangling Instance-level Moving Object Segmentation

arXiv:2606. 12826v1 Announce Type: cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking.

arXiv AI
Jul 13

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

arXiv:2607. 09114v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination variations, fast motion, and complex backgrounds when relying solely on visible light videos.

By Peipei Zhu, Yueqing Niu, Lin Zhu, Guanchong Niu, Yang Yu, Zheng Li
Hugging Face Trending Papers
Jul 13

Temporal Feature Distillation for Label-Efficient Precise Event Spotting in Sports Videos

Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.

arXiv Computer Vision
Sep 16

Hyper-RED: Scalable Event Pre-training via Semantic Hypergraph Distillation

Hyper-RED introduces a scalable image-to-event pretraining framework that transfers high‑order semantic structures via hypergraphs, avoiding rigid pixel‑wise alignment. By constructing image, event, and cross‑modal hypergraphs and applying a hypergraph relational distillation loss, the method preserves local relational consistency and event‑specific characteristics while inheriting image‑derived semantic organization. Experiments across five event datasets show consistent scaling from ViT‑S to ViT‑L and state‑of‑the‑art performance.

By Meisen Wang, Zhiqiang Tian, Wei Bao, Chengjie Wang, Shaoyi Du, Siqi Li
arXiv AI
Sep 15

EventVL: Understand Event Streams via Multimodal Large Language Model

EventVL introduces the first generative event-based multimodal large language model (MLLM) designed for explicit semantic understanding of event streams. The framework leverages a newly annotated dataset of nearly 1.4 million event–image/video–text pairs and incorporates an Event Spatiotemporal Representation to capture comprehensive event information, along with Dynamic Semantic Alignment to refine sparse semantic spaces. Experiments demonstrate that EventVL outperforms existing MLLM baselines in event captioning and scene description generation tasks, advancing the field of event vision.

By Pengteng Li, Yunfan Lu, Pinghao Song, Wuyang Li, Huizai Yao, Hui Xiong