Vinson·Li

Index

On multimodal

  1. GPT-4o hears you laugh

    OpenAI and Google both showed real-time multimodal assistants this week. Speech as a native modality changes what a conversation with a model is, and latency turns out to be the feature.

    2 min
  2. Natively multimodal

    Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.

    2 min
  3. GPT-4 reads the picture

    OpenAI's GPT-4 scores near the top of the bar exam and can explain a joke in a photo. Describing a scene is useful, but I'm still unsure how far that gets it toward understanding the physical world.

    2 min
  4. One architecture, any input

    DeepMind's Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.

    2 min
  5. Pictures and words in the same space

    OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.

    2 min
  6. Faces help you hear

    Google's Looking to Listen separates one voice from a crowd by watching the speaker's face. Modalities work better when they explain each other.

    2 min