Fake Drake
An AI-generated song imitating Drake and The Weeknd got millions of plays before it was pulled. Voice is identity, and the industry needs consent and attribution systems, fast.
A song called “Heart on My Sleeve,” credited to an anonymous account called ghostwriter977, went viral over the last two weeks. It uses AI-generated vocals that sound like Drake and The Weeknd. It got millions of plays and views across streaming services and social platforms before Universal Music Group, which represents both artists, had it taken down last week. UMG’s statement asked which side of history the industry wants to be on: “the side of artists, fans and human creative expression, or on the side of deep fakes, fraud and denying artists their due compensation.”
I work on YouTube Music. I’m not going to talk about anything specific to how any platform handled this. I’ll write about what I think the problem is.
Technically, this isn’t new. Voice conversion models have been improving for years, and open-source tools let anyone train a model on an artist’s vocals and convert their own singing into that voice. What’s new is that the result was good enough, and the song catchy enough, that millions of people listened because they liked it, not as a novelty.
A voice is different from a style. Stable Diffusion imitating an illustrator’s style raises real questions, but style has always been somewhat shared in art. Artists influence each other constantly. A voice is closer to a face. It identifies a specific person. When you hear Drake’s voice, you assume Drake chose to say those words. Putting words in someone’s voice without consent is impersonation, and it affects the artist’s reputation and their business directly.
What I think is needed, and fast:
A consent layer. Artists should be able to say whether their voice can be used for AI generation, by whom, and on what terms. Some artists will say no. Some will say yes and want a share. Grimes said on Sunday she’d split royalties 50/50 on any successful AI song using her voice. That’s the start of a business model.
Attribution. Generated content should be labeled, with some provenance attached, so listeners and platforms know what they’re hearing. That’s technically hard for audio, which is easily re-recorded, but not impossible.
Detection, as a backstop. Platforms will need to identify voice clones of known artists at upload time, the way they identify copyrighted recordings today.
None of this is simple, and none of it will be perfect. But I don’t think “ban it” works, since the tools are open, or that “anything goes” is acceptable. The industry dealt with sampling in the 80s and 90s, and ended up with licensing that let sampling become a normal part of music. Voice cloning will probably need something similar, and the platforms, labels and artists who build it first will set the terms for everyone.