Vinson·Li

Essay No. 126

Describe the music you want

"Sad songs for a rainy drive" beats any genre taxonomy. Natural language is becoming the way people ask for music, and it changes what a music recommender has to understand.


Over the last year, several music services have started letting people describe a playlist in their own words and generate it. Spotify rolled out AI Playlist in the US in September, after testing it in the UK and Australia. Other services are testing similar things. You type something like “upbeat songs for cleaning the house on a Sunday morning” or ”90s R&B for a dinner party that doesn’t get too loud,” and you get a playlist.

I work on YouTube Music; the examples here are from public products.

I’ve thought about music discovery a lot since I started this job, and I think this is one of the biggest changes in how people will ask for music since streaming began.

The way people have navigated music for decades is by genre, artist and chart. Those are useful, but they’re not how people think about what they want to hear. People think in moods, activities, memories and situations. “Something to focus to that isn’t boring.” “Songs that sound like the summer I was 19.” “Music my parents and my kids will both tolerate on a road trip.” None of those map to a genre. Until recently, the only way to serve them was an editor making a playlist with that exact name and hoping enough people wanted the same thing.

What’s changed is that we now have models that connect language and audio in the same space, like MuLan did for MusicLM, and language models that can interpret a messy request, break it into parts, and reason about what fits. So a request can be decomposed into several pieces: mood, tempo, era, energy arc, context. Then candidates can be retrieved by matching both the description and the listener’s own taste, and ordered into a sequence that holds together.

Getting from a request to a playlist leaves three problems to solve.

Interpretation. “Rainy drive” means something different to different people. The system has to make a reasonable guess and let you steer: more like this, less like that, no country.

Personalization. Two people typing the same words want different playlists, because their taste and history differ. A system that only understands the words, and not the person, will produce a generic answer.

Sequencing. As I wrote in 2021, a playlist is a sequence, not a pile. The request describes an experience over time, and the order matters as much as the selection.

I think describing what you want will become one of the main ways people find music, alongside radio-style streams and their own libraries. It turns the recommender from a system that guesses silently into one you talk to, which is a big shift in how the product feels.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…