arXiv Computer Vision By Rainer Lienhart, Daniel Kienzle, Shin'ichi Satoh, Anastasiia Bilinska

Event Detection in Table Tennis Videos using 2D Keypoints

Read the original on arXiv Computer Vision →

The paper introduces EventNet, a two‑stage pipeline that uses 2D keypoints of players, table corners, and the ball to detect key events in table tennis videos. First, a keypoint transformer condenses the pose and ball information into a robust representation; second, a transformer encoder predicts how close each frame is to the next and previous ball‑racket contact using a novel temporal cosine‑like target signal. Experiments on Latte‑MV and TTHQ datasets show high accuracy, with an F1 score of 91.16% and a mean frame deviation of 0.42 on Latte‑MV, and 73.08% / 1.16 on TTHQ.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 13

Temporal Feature Distillation for Label-Efficient Precise Event Spotting in Sports Videos

Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries.