BoT-Feedback: Grounding Multimodal Reasoning in Biomechanical Evidence for Explainable Human Action Feedback
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
MyoMechanix is a multimodal dataset and framework for action quality assessment that incorporates muscle activity and other physiological signals alongside visual data. It contains over 7,500 samples of 20 weight‑loaded actions from 38 subjects, with synchronized RGB video, 3D pose, sEMG, and additional signals. The accompanying Fitness Knowledge Graph structures expert annotations into relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment through the CUBIST engine. The project also introduces MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task, demonstrating that multimodal sensing and structured representations improve performance, interpretability, and error attribution.
DrGait is a training‑free framework that transforms Vision‑Language Models into clinical planners for gait analysis. It separates semantic reasoning from geometric perception using a Triage‑Verification‑Synthesis workflow, where hypotheses are generated, verified with deterministic biomechanical tools, and refined in a closed‑loop. This approach reduces hallucinations and produces transparent, audit‑ready clinical reports with competitive diagnostic accuracy.
arXiv:2607. 16322v1 Announce Type: cross Abstract: Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise.
arXiv:2609.39378v1 Announce Type: new Abstract: Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolvin...
PhysVista is a new benchmark that evaluates physical intelligence in Vision‑Language Models (VLMs) by integrating perception, reasoning, and plausibility assessment into a closed cognitive loop. It distinguishes between event‑level and scale‑level reasoning and tests models on both real‑world and AI‑generated videos to provide a holistic, fine‑grained analysis of physical understanding. Experiments show significant gaps in VLMs’ physical reasoning and plausibility assessment, underscoring the need for more principled, physically grounded multimodal designs.
PhysioAI introduces a clinical knowledge‑guided framework that injects structured physiotherapy knowledge into skeleton‑based action recognition models. By combining graph‑based spatiotemporal modeling with semantic anchors derived from a Clinical Knowledge Dictionary encoded via a frozen CLIP model, PhysioAI improves training of skeleton representations while requiring only skeleton inputs at inference. In subject‑disjoint evaluations, it outperforms existing methods on KiMoRe, Hard‑67, and UI‑PRMD benchmarks, achieving up to 99.03% accuracy on KiMoRe overall.