On language
- DeepSeek and the price of intelligence
A Chinese lab released a reasoning model that matches OpenAI's o1 and published how it did it. Nvidia lost about $600 billion in a day. Where the efficiency actually came from.
2 min reads likes comments - Describe the music you want
"Sad songs for a rainy drive" beats any genre taxonomy. Natural language is becoming the way people ask for music, and it changes what a music recommender has to understand.
2 min reads likes comments - GPT-4 reads the picture
OpenAI's GPT-4 scores near the top of the bar exam and can explain a joke in a photo. Describing a scene is useful, but I'm still unsure how far that gets it toward understanding the physical world.
2 min reads likes comments - ChatGPT is a product, not a model
A million people signed up in five days. The underlying model isn't new. What's new is the interface and the training to follow instructions, and that changes every software team.
2 min reads likes comments - We've been undertraining
DeepMind's Chinchilla paper says large language models have far too many parameters for the data they see. The fix is more data, and that raises a question about where it comes from.
2 min reads likes comments - Copilot finished my function
GitHub's Copilot preview suggests whole functions inside the editor. Two days of using it on side projects, and what changes when the IDE becomes a conversation.
2 min reads likes comments - Pictures and words in the same space
OpenAI's CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.
2 min reads likes comments - GPT-3 wrote a React component from a sentence
The demos flooding Twitter this week come from the same next-word objective as GPT-2, scaled a hundred times. Few-shot learning appeared on its own. Coding changes first.
2 min reads likes comments - "Too dangerous to release"
OpenAI trained a much bigger language model and is holding back the full version. The unicorn story is impressive. What's actually dangerous is cheap, plausible text at scale.
2 min reads likes comments - BERT reads both directions at once
Google's BERT beat almost every language benchmark by predicting hidden words using context on both sides. It's a representation model, which is a different thing from a text generator.
2 min reads likes comments - Read everything first, specialize later
OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.
2 min reads likes comments - No recurrence, no convolution
A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.
2 min reads likes comments - Google Translate got better overnight, and Chinese speakers noticed first
Neural machine translation replaced phrase tables, starting with Chinese to English. Notes from someone who reads both, and why the zero-shot result is the bigger story.
2 min reads likes comments - Chatbots are this year's apps. Tay says slow down
Facebook opened Messenger to bots at F8, two weeks after Microsoft's Tay learned racism from Twitter in a day. Why I'm telling clients to wait.
2 min reads likes comments - A computer captioned a photo. It didn't see the photo
Google and Stanford both have networks that write sentences about images. How they work, and what they're actually learning.
2 min reads likes comments