arXiv AI By Yuto Shibata, Yusuke Oumi, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa

BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 7

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, an encoder‑decoder framework that estimates 3D poses of multiple people using only acoustic signals. It addresses challenges such as overlapping acoustic signatures and inter‑person reflections by employing a multi‑scale acoustic encoder and a temporal pose decoder with attention. The authors created the 6‑hour Acoustic Multi‑person Pose (AMP) dataset and show that SoundMHPE outperforms baseline models.

By Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa
Hugging Face Trending Papers
Sep 4

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, the first system to estimate multi‑person 3D poses using only acoustic signals. It tackles challenges such as overlapping acoustic signatures and inter‑person reflections by employing an Acoustic Multi‑scale Encoder and a Temporal Pose Decoder with attention. The authors built a 6‑hour Acoustic Multi‑person Pose dataset and show that SoundMHPE outperforms baseline models.

arXiv AI
2d ago

Supervising Sound Localization by In-the-wild Egomotion

The paper introduces a method for learning binaural sound localization by using egomotion as a supervisory signal. By tracking how a camera’s direction changes relative to a sound source during a video, the authors train an audio model to predict sound directions that align with visual estimates of camera motion derived from multi‑view geometry. They evaluate this approach on a newly proposed dataset of real‑world audio‑visual videos with egomotion, demonstrating that the model can learn from real data and perform well on sound localization tasks.

By Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens
arXiv AI
Sep 24

SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

SsgCaps is a publicly available dataset of human-engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action-based typology, enabling extensive yet plausible sampling. A comparative quantitative analysis shows only small differences between the open and private versions, supporting the recommendation of the open version for benchmarking sound scene generation algorithms.

By Modan Tailleur (LS2N), Junwon Lee (LS2N), Laurie M Heller (LS2N), Mathieu Lagrange (LS2N), Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto