arXiv Machine Learning

Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

arXiv:2607. 18303v1 Announce Type: cross Abstract: Identifying which string produces a given pitch in monophonic electric guitar audio is a fundamental classification challenge: a single pitch can often be produced on multiple strings at different fret positions, with timbral differences that prior listening studies confirm are largely imperceptible to untrained humans.

arXiv Machine Learning
Sep 11

TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

TART is a modular four‑stage pipeline that transcribes guitar audio into tablature, addressing key limitations of existing systems such as missing expressive techniques, incorrect string‑fret assignments, and poor performance on noisy recordings. The stages include an audio‑to‑MIDI model, an expressive technique classifier, an audio‑conditioned T5 encoder‑decoder for string‑fret mapping, and an automated tablature generator. In zero‑shot evaluations on GuitarSet, EGDB, and noisy variants, TART outperforms prior baselines with significant gains in audio‑to‑MIDI, string‑fret, and end‑to‑end tablature metrics.

By Akshaj Gupta, Hwi Joo Park, Andrea Guzman, Shamak Gowda, Samhita Konduri, Jiachen Lian, Robin Netzorg, Gopala Anumanchipalli
arXiv Machine Learning
Sep 17

Instrument Classification of Solo Sheet Music Images

This paper investigates instrument classification using solo sheet music images rather than audio. It converts images into sequences of musical words via bootleg score representation and treats the task as text classification, training AWD‑LSTM, GPT‑2, and RoBERTa models on IMSLP data for eight instruments. Pretraining on unlabeled data and fine‑tuning improves RoBERTa’s accuracy from 34.5% to 42.9%, and two proposed data‑augmentation methods raise accuracy by an additional 15%.

By Kevin Ji, Daniel Yang, TJ Tsai
arXiv AI
Sep 2

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.

By Utsab Ghosh, Roshni Chakraborty
arXiv AI
Sep 2

TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data

TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.

By Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
arXiv Machine Learning
1d ago

LAST: Looped Audio Spectrogram Transformer

The paper introduces LAST, a Looped Audio Spectrogram Transformer that processes all tokens once and then reuses the same blocks to refine only the class token over fixed audio features, making subsequent passes inexpensive. On AudioSet, a ten‑pass LAST outperforms a twelve‑layer sequential transformer by 2.1% relative mean average precision while using 49.4% fewer parameters, 42% fewer MACs, and achieving 9.8% higher throughput. Increasing the pass count from two to ten improves accuracy with only a 1.2% increase in computation, and the model shows enhanced robustness to temporal masking and other auditory augmentations across music, environmental, and event sound classification tasks.

By Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty