Return
Generating coherent and efficient melodic continuations with deterministic discrete diffusion
DOI:10.1080/00051144.2025.2591993.png)
Abstract
En 中文
Melody continuation aims to automatically generate coherent melodic sequences conditioned on an initial musical motif, presenting significant challenges due to the intricate balance required between structural consistency, rhythmic coherence, and computational efficiency. This paper introduces Deterministic Discrete Diffusion for Melody Continuation (D3MC), a novel framework leveraging discrete diffusion processes with a deterministic mask-based sampling strategy. Unlike conventional autoregressive models suffering from slow sequential inference, or traditional diffusion models susceptible to rhythmic inconsistency, D3MC integrates a progressive timestep sampling schedule and a predominantly mask-driven noise injection strategy. This combination systematically preserves rhythmic and harmonic structures while enabling parallel prediction of melodic events. Evaluations conducted on the Lakh MIDI and MAESTRO datasets demonstrate superior performance of D3MC compared to leading baselines. Specifically, our model achieves a Pitch Contour Dynamic Time Warping (DTW) score of 13.1 (outperforming standard diffusion methods by approximately 10.9%), a Rhythm KL divergence of 0.44, and a Harmony compatibility score of 0.79, all achieved with an inference speed of merely 0.8 seconds per sequence-representing up to an 8.9-fold efficiency gain over autoregressive counterparts.
Keywords:
Discrete diffusion models
deterministic sampling
melody continuation
symbolic music generation
sequence-toarning-Sequence le

