arXiv AI

RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition

RAST is a resolution‑aware transfer framework that addresses the mismatch between high‑resolution (HR) audio available during training and low‑resolution (LR) audio used at inference for human activity recognition. By compressing HR teacher representations while preserving token‑level information and neighborhood structure, RAST performs localized HR‑LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR‑only training and direct teacher‑transfer baselines, improving LR‑only recognition by up to approximately 7.8% while requiring only LR audio at inference.

arXiv Machine Learning
Aug 28

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

HALO is a heterogeneity‑aware, language‑aligned foundation model for inertial measurement unit (IMU) based human activity recognition. It uses a two‑stage training process: first, a self‑supervised encoder learns to handle diverse sensor configurations and natural‑language sensor descriptions; second, the encoder is aligned with text embeddings through synonym‑aware contrastive learning, enabling open‑set recognition via cosine similarity. Trained on ten public HAR datasets, HALO outperforms five state‑of‑the‑art baselines across eight metrics while using only ~35 M parameters, and improves zero‑shot open‑set accuracy by 13.7 percentage points over 87 training labels.

By Zihan Ding, Liyu Zhang, Xiaomin Ouyang
arXiv AI
3d ago

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
arXiv Machine Learning
Jul 24

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

arXiv:2602. 10230v2 Announce Type: replace Abstract: Audio language models process input audio into rich frame-level representations, but the standard approach to temporal localization generates timestamps as sequences of text tokens, which discards the frame-level representations in favor of autoregressive decoding.

By Joseph An, Phillip Keung, Jiaqi Wang, Orevaoghene Ahia, Noah A. Smith
arXiv AI
Sep 1

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha