arXiv AI By Ji Hwan Park, Gautham Krishna Gudur, Yufei Shen, Dawei Liang, Edison Thomaz

RAST: Resolution-Aware Privileged Structure Transfer for Low-Resolution Audio Activity Recognition

Read the original on arXiv AI →

RAST is a resolution‑aware transfer framework that addresses the mismatch between high‑resolution (HR) audio available during training and low‑resolution (LR) audio used at inference for human activity recognition. By compressing HR teacher representations while preserving token‑level information and neighborhood structure, RAST performs localized HR‑LR alignment. Experiments on the SAMoSA and AudioIMU datasets show that RAST consistently outperforms LR‑only training and direct teacher‑transfer baselines, improving LR‑only recognition by up to approximately 7.8% while requiring only LR audio at inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition

HALO is a heterogeneity‑aware, language‑aligned foundation model for inertial measurement unit (IMU) based human activity recognition. It uses a two‑stage training process: first, a self‑supervised encoder learns to handle diverse sensor configurations and natural‑language sensor descriptions; second, the encoder is aligned with text embeddings through synonym‑aware contrastive learning, enabling open‑set recognition via cosine similarity. Trained on ten public HAR datasets, HALO outperforms five state‑of‑the‑art baselines across eight metrics while using only ~35 M parameters, and improves zero‑shot open‑set accuracy by 13.7 percentage points over 87 training labels.

By Zihan Ding, Liyu Zhang, Xiaomin Ouyang
arXiv AI
3d ago

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu