LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
arXiv:2607. 05769v1 Announce Type: cross Abstract: We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music.
Optical Music Recognition (OMR) has seen major progress in model design, with end-to-end methods now capable of recognising notation at all levels of complexity. However, the impact of this progress has been limited by the visual domains of available training datasets, which are largely born-digital.
arXiv:2607. 05769v1 Announce Type: cross Abstract: We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music.
arXiv:2609.05662v1 Announce Type: cross Abstract: Full-page end-to-end Optical Music Recognition seeks to transcribe entire music pages directly into symbolic notation, avoiding the limitations of tr...
Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form scores, multi-part score transcription remains underexplored, largely due to the absence of a suitable dataset.
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.
arXiv:2512. 02652v2 Announce Type: replace-cross Abstract: Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language.
arXiv:2607. 00777v1 Announce Type: cross Abstract: Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present.
arXiv:2601.11262v2 Announce Type: replace-cross Abstract: Music Cover Retrieval, also known as Version Identification, aims to recognize distinct renditions of the same underlying musical work, a tas...
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often...
ONOTE is a unified framework that treats music as a scientifically structured domain of measurable cross-representation correspondences, focusing on omnimodal notation processing centered on sheet music. It introduces a test-only benchmark drawing from diverse musical sources—including staff, Jianpu, and tablature—across varied genres, instruments, and structural conditions, with aligned multimodal derivatives. The framework supports four complementary tasks—score understanding, notation conversion, audio transcription, and symbolic generation—while constructing a provenance-bearing proposition hypergraph from external music-theory materials for evidence retrieval and deterministic validity checks.
This paper investigates instrument classification using solo sheet music images rather than audio. It converts images into sequences of musical words via bootleg score representation and treats the task as text classification, training AWD‑LSTM, GPT‑2, and RoBERTa models on IMSLP data for eight instruments. Pretraining on unlabeled data and fine‑tuning improves RoBERTa’s accuracy from 34.5% to 42.9%, and two proposed data‑augmentation methods raise accuracy by an additional 15%.
arXiv:2607. 08168v1 Announce Type: cross Abstract: Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes.