TART is a modular four‑stage pipeline that transcribes guitar audio into tablature, addressing key limitations of existing systems such as missing expressive techniques, incorrect string‑fret assignments, and poor performance on noisy recordings. The stages include an audio‑to‑MIDI model, an expressive technique classifier, an audio‑conditioned T5 encoder‑decoder for string‑fret mapping, and an automated tablature generator. In zero‑shot evaluations on GuitarSet, EGDB, and noisy variants, TART outperforms prior baselines with significant gains in audio‑to‑MIDI, string‑fret, and end‑to‑end tablature metrics.
By Akshaj Gupta, Hwi Joo Park, Andrea Guzman, Shamak Gowda, Samhita Konduri, Jiachen Lian, Robin Netzorg, Gopala Anumanchipalli
arXiv:2607. 08863v1 Announce Type: cross Abstract: We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio.
By Oliverio Bombicci Pontelli, Iran R. Roman
arXiv:2608. 06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored.
By Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju
arXiv:2607. 01974v1 Announce Type: cross Abstract: This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings according to the Broad Sound Taxonomy (BST).
By Beile Ning, Jiayi Yu, Zitong Wang, Yufei Hu, Wenjun Xu, Yuanhang Qian, Zhongxin Bai, Gongping Huang
This paper investigates instrument classification using solo sheet music images rather than audio. It converts images into sequences of musical words via bootleg score representation and treats the task as text classification, training AWD‑LSTM, GPT‑2, and RoBERTa models on IMSLP data for eight instruments. Pretraining on unlabeled data and fine‑tuning improves RoBERTa’s accuracy from 34.5% to 42.9%, and two proposed data‑augmentation methods raise accuracy by an additional 15%.
By Kevin Ji, Daniel Yang, TJ Tsai
MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.
By Utsab Ghosh, Roshni Chakraborty
TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.
By Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
arXiv:2606. 31338v1 Announce Type: cross Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts.
By Yujun Lee, Joonhyeok Shin, Hyoeun Kim, Kyuhong Shim
The paper introduces LAST, a Looped Audio Spectrogram Transformer that processes all tokens once and then reuses the same blocks to refine only the class token over fixed audio features, making subsequent passes inexpensive. On AudioSet, a ten‑pass LAST outperforms a twelve‑layer sequential transformer by 2.1% relative mean average precision while using 49.4% fewer parameters, 42% fewer MACs, and achieving 9.8% higher throughput. Increasing the pass count from two to ten improves accuracy with only a 1.2% increase in computation, and the model shows enhanced robustness to temporal masking and other auditory augmentations across music, environmental, and event sound classification tasks.
By Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty
arXiv:2607. 03806v1 Announce Type: cross Abstract: Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood.
By H\'ector Martel, Joe Hennessy-Priest, Taemin Cho
arXiv:2512. 13998v3 Announce Type: replace-cross Abstract: Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corpora.
By Qilin Li, C. L. Philip Chen, Tong Zhang
Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successfu...