arXiv Computer Vision

Field Converter: Geometry-Initialized Temporal Residual Refinement for World-Grounded Player Pose Estimation from Soccer Broadcasts

arXiv Computer Vision
Sep 3

A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames

The paper introduces a top‑down framework for accurately locating athletes in metric world coordinates using a single calibrated broadcast frame. It presents three main contributions: a Boundary‑Aware Adaptive Tiling method that expands tile boundaries to avoid splitting athletes across tiles, a specialized two‑keypoint estimator based on RTMPose‑X for pelvis and ground projection points, and a deterministic lift of 2D projections into 3D world coordinates via camera‑calibrated ray casting. The approach achieves a LocSim score of 97.44 and an mAP of 0.9128, surpassing the baseline by over 21 % on a public test set.

By Thanh-Khoi Nguyen, Hoang-Phuc Nguyen, Linh-Huynh, Minh-Triet Tran
arXiv Computer Vision
Sep 18

Monocular Visual Odometry without Calibration or Test-time Optimization

The paper introduces CalfVO, a monocular visual odometry system that operates without camera intrinsics, test‑time optimization, bundle adjustment, or loop closure. Using a transformer, it predicts relative poses with separate rotation and translation confidences over overlapping image windows, then aggregates these predictions via a confidence‑weighted module to produce a single trajectory. CalfVO achieves the highest accuracy among calibration‑free methods across five benchmarks and runs at 53 FPS, outperforming all baselines.

By Vladimir Yugay, Duy-Kien Nguyen, Theo Gevers, Cees G. M. Snoek, Martin R. Oswald
arXiv Computer Vision
1d ago

Event Detection in Table Tennis Videos using 2D Keypoints

The paper introduces EventNet, a two‑stage pipeline that uses 2D keypoints of players, table corners, and the ball to detect key events in table tennis videos. First, a keypoint transformer condenses the pose and ball information into a robust representation; second, a transformer encoder predicts how close each frame is to the next and previous ball‑racket contact using a novel temporal cosine‑like target signal. Experiments on Latte‑MV and TTHQ datasets show high accuracy, with an F1 score of 91.16% and a mean frame deviation of 0.42 on Latte‑MV, and 73.08% / 1.16 on TTHQ.

By Rainer Lienhart, Daniel Kienzle, Shin'ichi Satoh, Anastasiia Bilinska
arXiv Computer Vision
Sep 3

MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion

MuyBridge is an on-device system that estimates an athlete’s segmental center of mass (CoM) trajectory from a single phone camera video stream. It combines a compact 2D pose network with a distilled monocular depth network, fusing their outputs through anatomical and physical priors to produce metric CoM estimates without requiring 3D or task‑specific supervision. On the AthletePose3D dataset, MuyBridge achieves 33–41 mm vertical CoM error and 2.3–6.6 % absolute‑relative range error, delivering CoM estimates at 63 FPS on an iPhone 15 with asynchronous depth updates.

By Aidan Bradshaw, Marco Giordano, David Rode, Andreas Habersack, Elif Basokur, Annika Kruse, Markus Tilp, Michele Magno, Peter Wolf, Luca Benini, Christoph Leitner