arXiv Computer Vision

Event Detection in Table Tennis Videos using 2D Keypoints

The paper introduces EventNet, a two‑stage pipeline that uses 2D keypoints of players, table corners, and the ball to detect key events in table tennis videos. First, a keypoint transformer condenses the pose and ball information into a robust representation; second, a transformer encoder predicts how close each frame is to the next and previous ball‑racket contact using a novel temporal cosine‑like target signal. Experiments on Latte‑MV and TTHQ datasets show high accuracy, with an F1 score of 91.16% and a mean frame deviation of 0.42 on Latte‑MV, and 73.08% / 1.16 on TTHQ.

Hugging Face Trending Papers
Jul 13

Temporal Feature Distillation for Label-Efficient Precise Event Spotting in Sports Videos

Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.

arXiv AI
Jun 10

Integrated Real-Time Motion Tracking and AI Analysis for Athletic Performance Optimization

arXiv:2606. 09842v1 Announce Type: cross Abstract: Applying Human Pose Estimation (HPE) in real world environments remains a challenging task, this paper explores and surveys real time HPE approaches and their limitations in sports analysis for individuals, alongside developing a practical lightweight prototype for real world testing and usage.

By Parth Agrawal, Ronit, Sagar Kumar, Aashish Bhambri
arXiv Computer Vision
Sep 16

EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset

EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.

By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa