arXiv AI

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

arXiv:2608. 08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored.

arXiv AI
Sep 18

Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

The paper investigates whether open‑source Vision‑Language Models (VLMs) can perform zero‑shot action quality assessment (AQA) on Olympic diving videos. Using the AQA‑7 benchmark, the authors propose a regression framework that combines VLM‑generated semantic reasoning, phase‑level sub‑scores, TF‑IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. While individual VLMs achieve moderate Spearman correlations (<0.32), the ensemble approach boosts performance to 0.67, demonstrating that VLM‑derived textual reasoning features are more informative than raw numerical sub‑scores for AQA. whyItMatters":"The study shows that VLMs can serve as explainable, semi‑automated tools for evaluating sports performance, potentially aiding expert judging in complex, subjective Olympic events."

By Henry O. Velesaca, David Freire-Obregon, Luigi Miranda, Abel Reyes-Angulo
arXiv Machine Learning
Aug 27

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

MyoMechanix is a multimodal dataset and framework for action quality assessment that incorporates muscle activity and other physiological signals alongside visual data. It contains over 7,500 samples of 20 weight‑loaded actions from 38 subjects, with synchronized RGB video, 3D pose, sEMG, and additional signals. The accompanying Fitness Knowledge Graph structures expert annotations into relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment through the CUBIST engine. The project also introduces MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task, demonstrating that multimodal sensing and structured representations improve performance, interpretability, and error attribution.

By Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
arXiv Machine Learning
Aug 27

Multimodal Injury Risk Prediction in Tennis

The paper introduces PART, a multimodal predictive framework for tennis that combines physiological, training, sleep, questionnaire, jump, and video data from nine collegiate players to assess overall wellness, injury risk, physical capability, and playing style. Using machine learning and deep learning, PART provides holistic athlete assessments and forecasts specific injury risks to body areas such as elbows and knees. Evaluation shows strong performance in predicting wellness and injury risk, with potential benefits for recreational players who often injure themselves due to poor technique.

By Francisco Erramuspe Alvarez, Shobharani Polasa, Weihao Qu, Jay Wang, Ling Zheng
Hugging Face Trending Papers
Jun 3

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granularity, undifferentiated difficulty, limited annotation quality, and pervasive answer ambiguity, leaving them unable to diagnose where current models fail.

Hugging Face Trending Papers
Jun 25

TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment

Existing facial expression quality assessment (FEQA) methods typically produce only a severity score, without explicitly communicating the observable facial motion evidence that supports the prediction. This limits interpretability and makes it difficult to inspect the basis of model outputs in Parkinson's disease assessment.

arXiv AI
Jun 8

OGA-AID: Clinician-in-the-loop AI Report Drafting Assistant for Multimodal Observational Gait Analysis in Post-Stroke Rehabilitation

arXiv:2604. 05360v2 Announce Type: replace-cross Abstract: Gait analysis is essential in post-stroke rehabilitation but remains time-intensive and cognitively demanding, especially when clinicians must integrate gait videos and motion-capture data into structured reports.

By Khoi T. N. Nguyen, Nghia D. Nguyen, Hui Yu Koh, Patrick W. H. Kwong, Karen Sui Geok Chua, Ananda Sidarta, Baosheng Yu