DeepMind introduces MusicLM, a text-to-music generation model
The model composed several minutes of coherent audio from a written prompt or a hummed melody, released as a research paper rather than a product.
- Models & capabilities
- Minor
Google Research published MusicLM, a model that generated music from a written text description, treating the task as a hierarchical sequence-to-sequence problem and producing audio at 24 kHz. Given a prompt such as “a calming violin melody backed by a distorted guitar riff,” it could sustain several minutes of audio that stayed thematically consistent with the description, and the paper reported that it outperformed earlier text-to-music systems on both audio quality and adherence to the prompt.
Alongside plain text prompts, the researchers demonstrated melody conditioning — transforming a whistled or hummed tune into a fuller arrangement in the style named by a caption — and a “story mode” that chained several prompts to steer how a piece developed over time. Google also released MusicCaps, a set of roughly 5,500 music clips paired with expert-written text descriptions, as a benchmark for future text-to-music work.
MusicLM was published as a research paper with an accompanying examples page, not shipped as a consumer product, and Google did not release the model’s weights, citing risks including the potential for misappropriating creators’ style. It sat in a similar position to DeepMind’s earlier image and speech research: a capability demonstration published well ahead of any product. Google later built the underlying research into Lyria, the music-generation model it began surfacing in public products from 2024.