Text to music is a representation problem
Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.
Google Research published MusicLM last week, a model that generates music from a text description, like “a calming violin melody backed by a distorted guitar riff,” at 24 kHz, for up to several minutes. The samples on the project page are good enough that I’ve been playing them to friends who don’t work in tech. The model hasn’t been released. I work on YouTube Music but wasn’t involved in MusicLM; I’m going by the public paper and samples.
What I find most interesting is how it handles the representation problem. Raw audio is enormous: 24,000 numbers per second. Modeling it directly, the way WaveNet did in 2016, is too slow for minutes of music, and it’s hard for a model to keep long-range structure, like a melody that returns or a key that holds, when it’s looking at individual samples.
MusicLM uses three kinds of tokens, built from three existing models.
Acoustic tokens come from SoundStream, a neural audio codec. It compresses audio into discrete codes at a much lower rate, and can decode them back into high-quality sound. These capture timbre, texture and recording quality, what it actually sounds like.
Semantic tokens come from w2v-BERT, a self-supervised audio model. They’re coarser and capture structure over longer time spans: rhythm, melody, harmony, the arc of a piece.
Text-music tokens come from MuLan, a model trained contrastively on music and text, like CLIP is for images and text. It puts a piece of music and its description in the same embedding space.
Generation happens in stages. From the MuLan embedding of the text, a Transformer generates semantic tokens, the long-term structure. From those, another stage generates the coarse acoustic tokens, and then the fine ones, which SoundStream decodes into audio. So the model decides what the music is doing before it decides exactly how it sounds.
That split between meaning and sound is the right design, I think, and it’s close to how musicians think. You write the chord progression and the melody, then decide the instrumentation and the mix. It also means each level can be learned from different data. The semantic and acoustic models are trained on audio alone, which is plentiful. Only MuLan needs paired music and text, and it learned from music videos and their text rather than from carefully labeled data.
Google says it won’t release the model for now, citing concerns about memorizing training data and about creative content. The paper reports that a small fraction of outputs closely matched training examples. For music, where copyright and artists’ voices and styles are everything, that’s a real problem and not a technicality. I suspect the next year will be less about whether these models can make music, and more about under what terms.