Vinson·Li

Index

On audio

  1. Veo 3 has sound, and that changes the medium

    Google's Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.

    2 min
  2. Text to music is a representation problem

    Google Research's MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.

    2 min
  3. Faces help you hear

    Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.

    2 min
  4. 16,000 samples a second

    DeepMind's WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it's so slow.

    2 min