arXiv Machine Learning By Oliverio Bombicci Pontelli, Iran R. Roman

Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations

Read the original on arXiv Machine Learning →

arXiv:2607. 08863v1 Announce Type: cross Abstract: We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

TART is a modular four‑stage pipeline that transcribes guitar audio into tablature, addressing key limitations of existing systems such as missing expressive techniques, incorrect string‑fret assignments, and poor performance on noisy recordings. The stages include an audio‑to‑MIDI model, an expressive technique classifier, an audio‑conditioned T5 encoder‑decoder for string‑fret mapping, and an automated tablature generator. In zero‑shot evaluations on GuitarSet, EGDB, and noisy variants, TART outperforms prior baselines with significant gains in audio‑to‑MIDI, string‑fret, and end‑to‑end tablature metrics.

By Akshaj Gupta, Hwi Joo Park, Andrea Guzman, Shamak Gowda, Samhita Konduri, Jiachen Lian, Robin Netzorg, Gopala Anumanchipalli
arXiv Machine Learning
Jul 22

Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

arXiv:2607. 18303v1 Announce Type: cross Abstract: Identifying which string produces a given pitch in monophonic electric guitar audio is a fundamental classification challenge: a single pitch can often be produced on multiple strings at different fret positions, with timbral differences that prior listening studies confirm are largely imperceptible to untrained humans.

By Aadi Garg
arXiv Machine Learning
Jul 28

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

arXiv:2607. 23395v1 Announce Type: cross Abstract: Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production.

By Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy, Tatiana Gabruseva
arXiv AI
Sep 7

Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

The paper introduces a lightweight technique to steer the pitch content of audio generated by the Stable Audio Open diffusion model. A small convolutional probe (~125k parameters) is trained to decode frame‑level pitch‑class activations from the model’s latent space using paired audio and MIDI data. During inference, the frozen probe acts as a differentiable loss, guiding generation toward a user‑specified pitch‑class sequence without retraining the base model, and improves melodic coherence by 2.4× over the unguided baseline.

By Yushi Ye, Wilson Zheng, Yongyi Zang