arXiv AI

BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds

arXiv AI
Sep 7

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, an encoder‑decoder framework that estimates 3D poses of multiple people using only acoustic signals. It addresses challenges such as overlapping acoustic signatures and inter‑person reflections by employing a multi‑scale acoustic encoder and a temporal pose decoder with attention. The authors created the 6‑hour Acoustic Multi‑person Pose (AMP) dataset and show that SoundMHPE outperforms baseline models.

By Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa
Hugging Face Trending Papers
Sep 4

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, the first system to estimate multi‑person 3D poses using only acoustic signals. It tackles challenges such as overlapping acoustic signatures and inter‑person reflections by employing an Acoustic Multi‑scale Encoder and a Temporal Pose Decoder with attention. The authors built a 6‑hour Acoustic Multi‑person Pose dataset and show that SoundMHPE outperforms baseline models.

arXiv AI
2d ago

Supervising Sound Localization by In-the-wild Egomotion

The paper introduces a method for learning binaural sound localization by using egomotion as a supervisory signal. By tracking how a camera’s direction changes relative to a sound source during a video, the authors train an audio model to predict sound directions that align with visual estimates of camera motion derived from multi‑view geometry. They evaluate this approach on a newly proposed dataset of real‑world audio‑visual videos with egomotion, demonstrating that the model can learn from real data and perform well on sound localization tasks.

By Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens
arXiv AI
Sep 24

SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

SsgCaps is a publicly available dataset of human-engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action-based typology, enabling extensive yet plausible sampling. A comparative quantitative analysis shows only small differences between the open and private versions, supporting the recommendation of the open version for benchmarking sound scene generation algorithms.

By Modan Tailleur (LS2N), Junwon Lee (LS2N), Laurie M Heller (LS2N), Mathieu Lagrange (LS2N), Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto
Hugging Face Trending Papers
Aug 5

Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen

Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consumption. While many studies and datasets in neuromorphic vision have focused on automotive and drone applications, human-centric daily-life scenarios remain largely underrepresented, despite their importance for developing and benchmarking event-based perception systems.