Vinson·Li

Essay No. 119

Music generation's GPT-3 moment

Suno v3 and Udio generate full songs with vocals and lyrics from a sentence, and some of them are good. What that means for artists, platforms and listeners.


Two releases in three weeks changed what I think is possible in music generation. Suno released v3 in late March, and on Wednesday Udio came out of stealth. Both take a short description, like “a melancholic indie folk song about leaving a small town, female vocals,” and produce a complete song, a couple of minutes long, with verses, a chorus, vocals singing coherent lyrics, and a mixed, mastered sound.

I’ve spent a lot of evenings this week generating songs. Most are forgettable, which is also true of most human-written songs. But maybe one in ten is something I’d listen to again, and a couple were genuinely good, with melodies that stuck in my head the next morning. A year ago, MusicLM’s instrumentals were impressive as research. This is a different category. It’s the GPT-3 moment for music: the point where the output stops being interesting only because a machine made it.

I work on YouTube Music. These are my reactions to the publicly available tools.

For listeners, I think the effect is smaller than people fear in the short term. Most listening is to songs people already know and artists they have relationships with. Music is identity and memory, and a generated song doesn’t have a person behind it to follow, see live or grow up with. I expect generated music to show up first where the artist never mattered much: background music for videos, games, ads, stores and workouts.

For artists, it’s complicated. On one hand, these tools are amazing sketchpads. A songwriter can hear a rough idea arranged and sung in seconds. On the other hand, nobody knows exactly what Suno and Udio were trained on, and both companies have been careful not to say. If they were trained on commercial recordings without licenses, which many people suspect, it’s the same consent problem I wrote about with Fake Drake last year and with face datasets in 2019. I’d expect lawsuits from the major labels this year.

For platforms, there’s a supply problem coming. If anyone can generate a thousand songs an hour, streaming services will be flooded with generated tracks, some uploaded to farm royalties. That will force decisions about labeling, about how royalties are split, and about how recommendations should treat generated content. Recommenders have been built on the assumption that content is scarce and costly to make. That assumption is going away.

I expect these tools to become part of how music is made, the way synthesizers, samplers and Auto-Tune did, each of which was controversial at first. The artists who use them well will make things nobody could make before. And the legal framework will catch up slowly, through licensing deals between the model companies and the rights holders, after a period of lawsuits.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…