<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Vinson Li</title><description>Essays on AI, video, and the craft of building things, from an engineer who reads too much Dostoevsky.</description><link>https://vinsonli.com/</link><item><title>What would it cost to make Breaking Bad with AI?</title><link>https://vinsonli.com/posts/what-would-it-cost-to-make-breaking-bad-with-ai/</link><guid isPermaLink="true">https://vinsonli.com/posts/what-would-it-cost-to-make-breaking-bad-with-ai/</guid><description>A line-by-line estimate for one 47-minute episode using this year&apos;s public video generation prices. Generating the footage is only part of the budget; performance and continuity remain much harder to get right.</description><pubDate>Fri, 11 Sep 2026 00:00:00 GMT</pubDate><category>video</category><category>economics</category><category>filmmaking</category><category>microdrama</category></item><item><title>Taste without a verifier</title><link>https://vinsonli.com/posts/taste-without-a-verifier/</link><guid isPermaLink="true">https://vinsonli.com/posts/taste-without-a-verifier/</guid><description>Coding agents somehow learned to build good-looking websites, although nothing checks whether a design is beautiful. What video editing can borrow from that: preference models, critics and loops.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>taste</category><category>agents</category><category>creative-tools</category><category>video</category></item><item><title>Learning like a baby: a plan for an embodied world model</title><link>https://vinsonli.com/posts/learning-like-a-baby/</link><guid isPermaLink="true">https://vinsonli.com/posts/learning-like-a-baby/</guid><description>Google DeepMind&apos;s Gemini Robotics ER 2 gives robots a better high-level brain. The part I still think nobody has built is the body-first learning underneath. Here&apos;s the research plan I&apos;d run.</description><pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate><category>world-models</category><category>embodiment</category><category>robotics</category><category>curiosity</category><category>research</category></item><item><title>From embeddings to IDs to what you watch next</title><link>https://vinsonli.com/posts/from-embeddings-to-ids-to-what-you-watch-next/</link><guid isPermaLink="true">https://vinsonli.com/posts/from-embeddings-to-ids-to-what-you-watch-next/</guid><description>Generative recommenders are going into production. How a video becomes an embedding, then a semantic ID, and how a Transformer uses those IDs to choose what you&apos;ll watch next.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>recommendation</category><category>semantic-ids</category><category>transformer</category><category>video</category></item><item><title>Change everything except the motion</title><link>https://vinsonli.com/posts/change-everything-except-the-motion/</link><guid isPermaLink="true">https://vinsonli.com/posts/change-everything-except-the-motion/</guid><description>At I/O this week, YouTube brought Gemini Omni to Shorts Remix: take a Short, change the characters, setting or style, and keep the motion and story. Why remix is the right first experience for video-to-video.</description><pubDate>Fri, 22 May 2026 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>youtube</category><category>remix</category><category>creators</category></item><item><title>Sora is shutting down. Generation isn&apos;t a product</title><link>https://vinsonli.com/posts/sora-is-shutting-down/</link><guid isPermaLink="true">https://vinsonli.com/posts/sora-is-shutting-down/</guid><description>OpenAI closed the Sora app this week, seven months after launching it as a social network, and the API ends in September. What went wrong, and what it says about where generative video belongs.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>product</category><category>social</category></item><item><title>A billion dollars for JEPA</title><link>https://vinsonli.com/posts/a-billion-dollars-for-jepa/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-billion-dollars-for-jepa/</guid><description>Yann LeCun&apos;s AMI Labs raised $1.03 billion to build world models on the JEPA framework. Revisiting what I wrote in 2022, and the one thing I still think the plan is missing.</description><pubDate>Fri, 13 Mar 2026 00:00:00 GMT</pubDate><category>world-models</category><category>jepa</category><category>startups</category><category>embodiment</category></item><item><title>The model plans the shots now</title><link>https://vinsonli.com/posts/the-model-plans-the-shots-now/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-model-plans-the-shots-now/</guid><description>ByteDance&apos;s Seedance 2.0 generates multi-shot sequences with references for characters, props and sound, and triggered cease-and-desist letters within a day. What&apos;s left for the editor, and for rights holders.</description><pubDate>Tue, 17 Feb 2026 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>filmmaking</category><category>copyright</category></item><item><title>A weekend with the World API</title><link>https://vinsonli.com/posts/a-weekend-with-the-world-api/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-weekend-with-the-world-api/</guid><description>World Labs opened an API for Marble. I spent the weekend building a small workbench for rendering its worlds as Gaussian splats and authoring camera paths through them. What broke: scale, coordinates and drift.</description><pubDate>Sun, 25 Jan 2026 00:00:00 GMT</pubDate><category>world-models</category><category>3d</category><category>gaussian-splatting</category><category>side-projects</category></item><item><title>What a video editing agent needs to get right</title><link>https://vinsonli.com/posts/video-editing-isnt-a-knapsack-problem/</link><guid isPermaLink="true">https://vinsonli.com/posts/video-editing-isnt-a-knapsack-problem/</guid><description>I built an agent that edits my social videos end to end using knowledge of my personal taste. What that asks of the timeline, the feedback loop and the product.</description><pubDate>Tue, 16 Dec 2025 00:00:00 GMT</pubDate><category>video</category><category>editing</category><category>agents</category><category>creative-tools</category></item><item><title>LeCun leaves, Marble ships</title><link>https://vinsonli.com/posts/lecun-leaves-marble-ships/</link><guid isPermaLink="true">https://vinsonli.com/posts/lecun-leaves-marble-ships/</guid><description>In one week World Labs launched Marble, Yann LeCun confirmed he&apos;s leaving Meta to start a world-model company, and Google shipped Gemini 3. The world-model bet went public from three directions.</description><pubDate>Sun, 23 Nov 2025 00:00:00 GMT</pubDate><category>world-models</category><category>jepa</category><category>3d</category><category>industry</category></item><item><title>Sora built a social network</title><link>https://vinsonli.com/posts/sora-built-a-social-network/</link><guid isPermaLink="true">https://vinsonli.com/posts/sora-built-a-social-network/</guid><description>OpenAI launched Sora 2 as a TikTok-style app where every video is generated. The model is impressive. I&apos;m less sure a feed of generations gives people a reason to keep watching.</description><pubDate>Mon, 06 Oct 2025 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>social</category><category>product</category></item><item><title>Consistent characters are the unlock</title><link>https://vinsonli.com/posts/consistent-characters-are-the-unlock/</link><guid isPermaLink="true">https://vinsonli.com/posts/consistent-characters-are-the-unlock/</guid><description>Google&apos;s new image model, known everywhere as Nano Banana, keeps a person looking like themselves across edits and scenes. For AI storytelling, identity preservation matters more than image quality.</description><pubDate>Sun, 14 Sep 2025 00:00:00 GMT</pubDate><category>generative-models</category><category>images</category><category>storytelling</category><category>consistency</category></item><item><title>Genie 3 remembers where you painted the wall</title><link>https://vinsonli.com/posts/genie-3-remembers-where-you-painted-the-wall/</link><guid isPermaLink="true">https://vinsonli.com/posts/genie-3-remembers-where-you-painted-the-wall/</guid><description>DeepMind&apos;s Genie 3 generates interactive worlds in real time at 720p that stay consistent for minutes. Persistence is the test for a world model, and it just got much better.</description><pubDate>Sat, 09 Aug 2025 00:00:00 GMT</pubDate><category>world-models</category><category>deepmind</category><category>video</category><category>interactive</category></item><item><title>Sixty episodes, one minute each</title><link>https://vinsonli.com/posts/sixty-episodes-one-minute-each/</link><guid isPermaLink="true">https://vinsonli.com/posts/sixty-episodes-one-minute-each/</guid><description>Microdramas are one of the fastest-growing forms of entertainment, and the format exists because of production cost. What happens when AI changes the cost?</description><pubDate>Sun, 27 Jul 2025 00:00:00 GMT</pubDate><category>video</category><category>microdrama</category><category>economics</category><category>generative-models</category></item><item><title>62 hours of robot data</title><link>https://vinsonli.com/posts/62-hours-of-robot-data/</link><guid isPermaLink="true">https://vinsonli.com/posts/62-hours-of-robot-data/</guid><description>Meta&apos;s V-JEPA 2 learns a world model from a million hours of video, then learns to plan robot actions from 62 hours of robot data. Promising, and still missing touch.</description><pubDate>Mon, 16 Jun 2025 00:00:00 GMT</pubDate><category>world-models</category><category>jepa</category><category>robotics</category><category>video</category></item><item><title>Veo 3 has sound, and that changes the medium</title><link>https://vinsonli.com/posts/veo-3-has-sound/</link><guid isPermaLink="true">https://vinsonli.com/posts/veo-3-has-sound/</guid><description>Google&apos;s Veo 3 generates video with synchronized dialogue, sound effects and ambient audio. Sound turns clips into scenes. Also at I/O: a language model that writes by denoising.</description><pubDate>Sat, 24 May 2025 00:00:00 GMT</pubDate><category>video</category><category>audio</category><category>generative-models</category><category>google</category></item><item><title>What a minute of AI video costs in Beijing vs. San Francisco</title><link>https://vinsonli.com/posts/what-a-minute-of-ai-video-costs/</link><guid isPermaLink="true">https://vinsonli.com/posts/what-a-minute-of-ai-video-costs/</guid><description>Real per-second prices for Veo 2, Kling 2.0 and Jimeng this month, and why the retake rate matters more than the sticker price.</description><pubDate>Thu, 24 Apr 2025 00:00:00 GMT</pubDate><category>video</category><category>economics</category><category>generative-models</category><category>china</category></item><item><title>Style is free now. Taste isn&apos;t</title><link>https://vinsonli.com/posts/style-is-free-now-taste-isnt/</link><guid isPermaLink="true">https://vinsonli.com/posts/style-is-free-now-taste-isnt/</guid><description>GPT-4o&apos;s image generation turned the internet into Studio Ghibli for a week. When every style is one prompt away, the scarce thing is judgment.</description><pubDate>Sat, 29 Mar 2025 00:00:00 GMT</pubDate><category>generative-models</category><category>taste</category><category>creative-tools</category><category>images</category></item><item><title>Helix has a fast brain and a slow brain. So does my golf swing</title><link>https://vinsonli.com/posts/helix-and-my-golf-swing/</link><guid isPermaLink="true">https://vinsonli.com/posts/helix-and-my-golf-swing/</guid><description>Figure&apos;s Helix splits robot control into a slow model that understands and a fast one that moves. A golf swing works the same way, which is why thinking during it ruins it.</description><pubDate>Mon, 24 Feb 2025 00:00:00 GMT</pubDate><category>robotics</category><category>control</category><category>golf</category><category>humanoids</category></item><item><title>DeepSeek and the price of intelligence</title><link>https://vinsonli.com/posts/deepseek-and-the-price-of-intelligence/</link><guid isPermaLink="true">https://vinsonli.com/posts/deepseek-and-the-price-of-intelligence/</guid><description>A Chinese lab released a reasoning model that matches OpenAI&apos;s o1 and published how it did it. Nvidia lost about $600 billion in a day. Where the efficiency actually came from.</description><pubDate>Tue, 28 Jan 2025 00:00:00 GMT</pubDate><category>language</category><category>reasoning</category><category>china</category><category>economics</category></item><item><title>Genie 2 and World Labs in the same week</title><link>https://vinsonli.com/posts/genie-2-and-world-labs-in-the-same-week/</link><guid isPermaLink="true">https://vinsonli.com/posts/genie-2-and-world-labs-in-the-same-week/</guid><description>DeepMind&apos;s Genie 2 turns one image into a playable 3D world, and Fei-Fei Li&apos;s World Labs turns one image into a 3D scene you can walk through. Worlds are the next medium after text, images and video.</description><pubDate>Sun, 08 Dec 2024 00:00:00 GMT</pubDate><category>world-models</category><category>3d</category><category>generative-models</category><category>deepmind</category></item><item><title>Describe the music you want</title><link>https://vinsonli.com/posts/describe-the-music-you-want/</link><guid isPermaLink="true">https://vinsonli.com/posts/describe-the-music-you-want/</guid><description>&quot;Sad songs for a rainy drive&quot; beats any genre taxonomy. Natural language is becoming the way people ask for music, and it changes what a music recommender has to understand.</description><pubDate>Tue, 19 Nov 2024 00:00:00 GMT</pubDate><category>music</category><category>recommendation</category><category>language</category><category>product</category></item><item><title>The physics Nobel went to neural networks</title><link>https://vinsonli.com/posts/the-physics-nobel-went-to-neural-networks/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-physics-nobel-went-to-neural-networks/</guid><description>Hopfield and Hinton won the physics prize, and Hassabis and Jumper shared chemistry for AlphaFold. A mechanical engineer&apos;s reaction to physics and AI converging.</description><pubDate>Sat, 12 Oct 2024 00:00:00 GMT</pubDate><category>ai</category><category>physics</category><category>science</category><category>history</category></item><item><title>AI&apos;s bottleneck is a power plant</title><link>https://vinsonli.com/posts/ais-bottleneck-is-a-power-plant/</link><guid isPermaLink="true">https://vinsonli.com/posts/ais-bottleneck-is-a-power-plant/</guid><description>Microsoft signed a 20-year deal to restart a reactor at Three Mile Island to power its data centers. The constraint on AI is moving from chips to electricity, and baseload nuclear is suddenly back.</description><pubDate>Tue, 24 Sep 2024 00:00:00 GMT</pubDate><category>energy</category><category>infrastructure</category><category>economics</category><category>nuclear</category></item><item><title>A neural net runs Doom</title><link>https://vinsonli.com/posts/a-neural-net-runs-doom/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-neural-net-runs-doom/</guid><description>Google Research&apos;s GameNGen simulates Doom in real time with a diffusion model: no game engine, just next-frame prediction conditioned on your inputs. Interactive video is a world model.</description><pubDate>Sat, 31 Aug 2024 00:00:00 GMT</pubDate><category>world-models</category><category>games</category><category>diffusion</category><category>video</category></item><item><title>Search plus verification</title><link>https://vinsonli.com/posts/search-plus-verification/</link><guid isPermaLink="true">https://vinsonli.com/posts/search-plus-verification/</guid><description>AlphaProof and AlphaGeometry 2 together solved four of six problems from this year&apos;s Math Olympiad, reaching silver-medal level. What makes a checkable answer so useful for training, and what happens when there isn&apos;t one.</description><pubDate>Mon, 29 Jul 2024 00:00:00 GMT</pubDate><category>reasoning</category><category>deepmind</category><category>reinforcement-learning</category><category>taste</category></item><item><title>Kling came from a short-video company, not a lab</title><link>https://vinsonli.com/posts/kling-came-from-a-short-video-company/</link><guid isPermaLink="true">https://vinsonli.com/posts/kling-came-from-a-short-video-company/</guid><description>Kuaishou, TikTok&apos;s main rival in China, released a video model that rivals Sora&apos;s samples, and ordinary users in China can already try it. Video models get built by whoever has the video.</description><pubDate>Wed, 12 Jun 2024 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>china</category><category>short-video</category></item><item><title>GPT-4o hears you laugh</title><link>https://vinsonli.com/posts/gpt-4o-hears-you-laugh/</link><guid isPermaLink="true">https://vinsonli.com/posts/gpt-4o-hears-you-laugh/</guid><description>OpenAI and Google both showed real-time multimodal assistants this week. Speech as a native modality changes what a conversation with a model is, and latency turns out to be the feature.</description><pubDate>Fri, 17 May 2024 00:00:00 GMT</pubDate><category>multimodal</category><category>voice</category><category>agents</category><category>product</category></item><item><title>Music generation&apos;s GPT-3 moment</title><link>https://vinsonli.com/posts/music-generations-gpt-3-moment/</link><guid isPermaLink="true">https://vinsonli.com/posts/music-generations-gpt-3-moment/</guid><description>Suno v3 and Udio generate full songs with vocals and lyrics from a sentence, and some of them are good. What that means for artists, platforms and listeners.</description><pubDate>Sun, 14 Apr 2024 00:00:00 GMT</pubDate><category>music</category><category>generative-models</category><category>creators</category></item><item><title>Genie learned to play from videos with no controls</title><link>https://vinsonli.com/posts/genie-learned-to-play-from-videos/</link><guid isPermaLink="true">https://vinsonli.com/posts/genie-learned-to-play-from-videos/</guid><description>DeepMind&apos;s Genie learned a controllable world model from 2D platformer videos with no action labels, by inferring eight latent actions on its own. The model learns the controls as well as the game.</description><pubDate>Tue, 05 Mar 2024 00:00:00 GMT</pubDate><category>world-models</category><category>video</category><category>deepmind</category><category>games</category></item><item><title>&quot;World simulator&quot; is doing a lot of work</title><link>https://vinsonli.com/posts/world-simulator-is-doing-a-lot-of-work/</link><guid isPermaLink="true">https://vinsonli.com/posts/world-simulator-is-doing-a-lot-of-work/</guid><description>OpenAI&apos;s Sora generates minute-long videos that look astonishing. Its technical report calls it a world simulator. The samples show both why that&apos;s tempting and why it isn&apos;t true yet.</description><pubDate>Mon, 19 Feb 2024 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>world-models</category><category>openai</category></item><item><title>Mobile ALOHA cooks shrimp for $32,000</title><link>https://vinsonli.com/posts/mobile-aloha-cooks-shrimp/</link><guid isPermaLink="true">https://vinsonli.com/posts/mobile-aloha-cooks-shrimp/</guid><description>Stanford&apos;s two-armed robot on a wheeled base learned to cook, wipe spills and call an elevator from about fifty demonstrations per task. Cheap hardware plus teleoperation is changing robot data.</description><pubDate>Tue, 09 Jan 2024 00:00:00 GMT</pubDate><category>robotics</category><category>imitation-learning</category><category>data</category></item><item><title>Natively multimodal</title><link>https://vinsonli.com/posts/natively-multimodal/</link><guid isPermaLink="true">https://vinsonli.com/posts/natively-multimodal/</guid><description>Google announced Gemini, trained from the start on text, images, audio and video together. Why training on mixed modalities from day one matters more than bolting vision onto a language model.</description><pubDate>Tue, 12 Dec 2023 00:00:00 GMT</pubDate><category>multimodal</category><category>gemini</category><category>google</category><category>architecture</category></item><item><title>Four-second clips can&apos;t make a movie</title><link>https://vinsonli.com/posts/four-second-clips-cant-make-a-movie/</link><guid isPermaLink="true">https://vinsonli.com/posts/four-second-clips-cant-make-a-movie/</guid><description>Pika 1.0, Stable Video Diffusion and Runway&apos;s latest make beautiful short clips. What&apos;s missing is the grammar of film: continuity, screen direction and eyelines.</description><pubDate>Tue, 28 Nov 2023 00:00:00 GMT</pubDate><category>video</category><category>generative-models</category><category>filmmaking</category><category>creative-tools</category></item><item><title>The model rewrites your prompt</title><link>https://vinsonli.com/posts/the-model-rewrites-your-prompt/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-model-rewrites-your-prompt/</guid><description>DALL·E 3 in ChatGPT doesn&apos;t use what you typed. It has a language model write a long, detailed prompt for you. Creative tools are turning into agents that interpret intent.</description><pubDate>Thu, 19 Oct 2023 00:00:00 GMT</pubDate><category>generative-models</category><category>creative-tools</category><category>agents</category><category>openai</category></item><item><title>Optimus sorts blocks. A toddler falls down a thousand times</title><link>https://vinsonli.com/posts/optimus-sorts-blocks-a-toddler-falls-down/</link><guid isPermaLink="true">https://vinsonli.com/posts/optimus-sorts-blocks-a-toddler-falls-down/</guid><description>Tesla&apos;s new video shows Optimus sorting colored blocks with a neural network trained end to end. Curated robot demos and messy child learning, and which one scales.</description><pubDate>Sat, 30 Sep 2023 00:00:00 GMT</pubDate><category>robotics</category><category>humanoids</category><category>learning</category><category>embodiment</category></item><item><title>Gaussian splatting killed my mesh nostalgia</title><link>https://vinsonli.com/posts/gaussian-splatting-killed-my-mesh-nostalgia/</link><guid isPermaLink="true">https://vinsonli.com/posts/gaussian-splatting-killed-my-mesh-nostalgia/</guid><description>3D Gaussian Splatting represents a scene as millions of fuzzy, colored blobs and renders it in real time at NeRF quality. It does this without a neural network.</description><pubDate>Thu, 24 Aug 2023 00:00:00 GMT</pubDate><category>3d</category><category>graphics</category><category>representation</category><category>gaussian-splatting</category></item><item><title>Predicting in representation space</title><link>https://vinsonli.com/posts/predicting-in-representation-space/</link><guid isPermaLink="true">https://vinsonli.com/posts/predicting-in-representation-space/</guid><description>Meta&apos;s I-JEPA is the first concrete result from LeCun&apos;s world model agenda. It learns image representations by predicting hidden regions in latent space, with no augmentations and no pixel reconstruction.</description><pubDate>Tue, 11 Jul 2023 00:00:00 GMT</pubDate><category>jepa</category><category>self-supervised-learning</category><category>representation</category><category>world-models</category></item><item><title>Vision Pro scans your face to make you</title><link>https://vinsonli.com/posts/vision-pro-scans-your-face-to-make-you/</link><guid isPermaLink="true">https://vinsonli.com/posts/vision-pro-scans-your-face-to-make-you/</guid><description>Apple&apos;s headset builds a 3D &apos;Persona&apos; of your face so you can appear on video calls while wearing it. Ten years after my thesis, 3D face capture ships inside a headset.</description><pubDate>Thu, 08 Jun 2023 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>vr</category><category>apple</category><category>avatars</category></item><item><title>Recommendation as next-token prediction</title><link>https://vinsonli.com/posts/recommendation-as-next-token-prediction/</link><guid isPermaLink="true">https://vinsonli.com/posts/recommendation-as-next-token-prediction/</guid><description>A new paper turns every item into a short code of semantic tokens, then has a Transformer generate the code of what you&apos;ll want next. The item vocabulary finally describes what things are.</description><pubDate>Sun, 21 May 2023 00:00:00 GMT</pubDate><category>recommendation</category><category>semantic-ids</category><category>transformer</category><category>representation</category></item><item><title>Fake Drake</title><link>https://vinsonli.com/posts/fake-drake/</link><guid isPermaLink="true">https://vinsonli.com/posts/fake-drake/</guid><description>An AI-generated song imitating Drake and The Weeknd got millions of plays before it was pulled. Voice is identity, and the industry needs consent and attribution systems, fast.</description><pubDate>Tue, 25 Apr 2023 00:00:00 GMT</pubDate><category>music</category><category>generative-models</category><category>voice</category><category>ethics</category></item><item><title>GPT-4 reads the picture</title><link>https://vinsonli.com/posts/gpt-4-reads-the-picture/</link><guid isPermaLink="true">https://vinsonli.com/posts/gpt-4-reads-the-picture/</guid><description>OpenAI&apos;s GPT-4 scores near the top of the bar exam and can explain a joke in a photo. Describing a scene is useful, but I&apos;m still unsure how far that gets it toward understanding the physical world.</description><pubDate>Fri, 17 Mar 2023 00:00:00 GMT</pubDate><category>language</category><category>multimodal</category><category>openai</category><category>gpt-4</category></item><item><title>Control beats prompts</title><link>https://vinsonli.com/posts/control-beats-prompts/</link><guid isPermaLink="true">https://vinsonli.com/posts/control-beats-prompts/</guid><description>ControlNet lets you steer Stable Diffusion with a pose skeleton, a depth map or an edge sketch. Creators want to set the structure directly, and this gives them a way to.</description><pubDate>Mon, 20 Feb 2023 00:00:00 GMT</pubDate><category>generative-models</category><category>diffusion</category><category>creative-tools</category></item><item><title>Text to music is a representation problem</title><link>https://vinsonli.com/posts/text-to-music-is-a-representation-problem/</link><guid isPermaLink="true">https://vinsonli.com/posts/text-to-music-is-a-representation-problem/</guid><description>Google Research&apos;s MusicLM generates music from text descriptions. The interesting part is its stack of tokens: one for meaning, one for sound, one shared between music and words.</description><pubDate>Tue, 31 Jan 2023 00:00:00 GMT</pubDate><category>music</category><category>generative-models</category><category>audio</category><category>representation</category></item><item><title>ChatGPT is a product, not a model</title><link>https://vinsonli.com/posts/chatgpt-is-a-product-not-a-model/</link><guid isPermaLink="true">https://vinsonli.com/posts/chatgpt-is-a-product-not-a-model/</guid><description>A million people signed up in five days. The underlying model isn&apos;t new. What&apos;s new is the interface and the training to follow instructions, and that changes every software team.</description><pubDate>Tue, 06 Dec 2022 00:00:00 GMT</pubDate><category>language</category><category>chatgpt</category><category>product</category><category>openai</category></item><item><title>Optimus walked on stage. Watch its hands</title><link>https://vinsonli.com/posts/optimus-walked-on-stage-watch-its-hands/</link><guid isPermaLink="true">https://vinsonli.com/posts/optimus-walked-on-stage-watch-its-hands/</guid><description>A year after the dancer in the suit, Tesla showed a real humanoid prototype. It walked slowly and waved. The legs get the attention. The hands are the harder and more important problem.</description><pubDate>Tue, 04 Oct 2022 00:00:00 GMT</pubDate><category>robotics</category><category>humanoids</category><category>tesla</category><category>hands</category></item><item><title>Text to video is next</title><link>https://vinsonli.com/posts/text-to-video-is-next/</link><guid isPermaLink="true">https://vinsonli.com/posts/text-to-video-is-next/</guid><description>Meta&apos;s Make-A-Video generates short clips from a sentence. They&apos;re five seconds long, low resolution and physically wrong in instructive ways. The hard parts are consistency, continuity and physics.</description><pubDate>Fri, 30 Sep 2022 00:00:00 GMT</pubDate><category>generative-models</category><category>video</category><category>diffusion</category></item><item><title>Stable Diffusion runs on my own computer</title><link>https://vinsonli.com/posts/stable-diffusion-runs-on-my-own-computer/</link><guid isPermaLink="true">https://vinsonli.com/posts/stable-diffusion-runs-on-my-own-computer/</guid><description>Stability AI released the weights of a text-to-image model anyone can run on a consumer GPU. Open models change who gets to build, and what gets built.</description><pubDate>Sat, 27 Aug 2022 00:00:00 GMT</pubDate><category>generative-models</category><category>diffusion</category><category>open-source</category></item><item><title>LeCun&apos;s path, read carefully</title><link>https://vinsonli.com/posts/lecuns-path-read-carefully/</link><guid isPermaLink="true">https://vinsonli.com/posts/lecuns-path-read-carefully/</guid><description>Yann LeCun&apos;s position paper argues that intelligence needs world models that predict in representation space, not pixels. Where I agree, and where I&apos;d push back: bodies and hands.</description><pubDate>Sun, 03 Jul 2022 00:00:00 GMT</pubDate><category>world-models</category><category>jepa</category><category>self-supervised-learning</category><category>research</category></item><item><title>It learned Minecraft by watching YouTube</title><link>https://vinsonli.com/posts/it-learned-minecraft-by-watching-youtube/</link><guid isPermaLink="true">https://vinsonli.com/posts/it-learned-minecraft-by-watching-youtube/</guid><description>OpenAI&apos;s VPT labeled 70,000 hours of Minecraft videos with the actions players took, using a small model trained on a little labeled data. Passive video became interaction data.</description><pubDate>Sun, 26 Jun 2022 00:00:00 GMT</pubDate><category>agents</category><category>video</category><category>learning</category><category>openai</category></item><item><title>One network, 604 tasks</title><link>https://vinsonli.com/posts/one-network-604-tasks/</link><guid isPermaLink="true">https://vinsonli.com/posts/one-network-604-tasks/</guid><description>DeepMind&apos;s Gato plays Atari, captions images, chats and stacks blocks with a real robot arm, all with the same weights. Turning actions into tokens is the interesting part.</description><pubDate>Mon, 16 May 2022 00:00:00 GMT</pubDate><category>agents</category><category>transformer</category><category>robotics</category><category>deepmind</category></item><item><title>Images are solved-ish. Video is where physics lives</title><link>https://vinsonli.com/posts/images-are-solved-ish-video-is-where-physics-lives/</link><guid isPermaLink="true">https://vinsonli.com/posts/images-are-solved-ish-video-is-where-physics-lives/</guid><description>DALL·E 2 generates images that look like real photos and paintings from a sentence. Why the jump to video is much harder than the jump from GANs to this.</description><pubDate>Tue, 12 Apr 2022 00:00:00 GMT</pubDate><category>generative-models</category><category>diffusion</category><category>video</category><category>physics</category></item><item><title>We&apos;ve been undertraining</title><link>https://vinsonli.com/posts/weve-been-undertraining/</link><guid isPermaLink="true">https://vinsonli.com/posts/weve-been-undertraining/</guid><description>DeepMind&apos;s Chinchilla paper says large language models have far too many parameters for the data they see. The fix is more data, and that raises a question about where it comes from.</description><pubDate>Thu, 31 Mar 2022 00:00:00 GMT</pubDate><category>language</category><category>scaling</category><category>deepmind</category><category>data</category></item><item><title>Meta&apos;s worst day was TikTok&apos;s best</title><link>https://vinsonli.com/posts/metas-worst-day-was-tiktoks-best/</link><guid isPermaLink="true">https://vinsonli.com/posts/metas-worst-day-was-tiktoks-best/</guid><description>Meta lost over $230 billion in market value in one day after Facebook&apos;s first ever drop in daily users. The interest graph beat the social graph.</description><pubDate>Sat, 05 Feb 2022 00:00:00 GMT</pubDate><category>recommendation</category><category>short-video</category><category>social</category><category>business</category></item><item><title>NeRF went from hours to seconds</title><link>https://vinsonli.com/posts/nerf-went-from-hours-to-seconds/</link><guid isPermaLink="true">https://vinsonli.com/posts/nerf-went-from-hours-to-seconds/</guid><description>Nvidia&apos;s Instant NGP trains a neural radiance field in seconds with a multiresolution hash table. The representation mattered more than the network.</description><pubDate>Mon, 24 Jan 2022 00:00:00 GMT</pubDate><category>3d</category><category>graphics</category><category>representation</category><category>nerf</category></item><item><title>Diffusion is going to eat GANs</title><link>https://vinsonli.com/posts/diffusion-is-going-to-eat-gans/</link><guid isPermaLink="true">https://vinsonli.com/posts/diffusion-is-going-to-eat-gans/</guid><description>Two papers this week, GLIDE and latent diffusion, make text-to-image generation with diffusion models look practical. What denoising actually learns, and why it beats the adversarial game.</description><pubDate>Tue, 28 Dec 2021 00:00:00 GMT</pubDate><category>generative-models</category><category>diffusion</category><category>images</category></item><item><title>Facebook deletes a billion faceprints</title><link>https://vinsonli.com/posts/facebook-deletes-a-billion-faceprints/</link><guid isPermaLink="true">https://vinsonli.com/posts/facebook-deletes-a-billion-faceprints/</guid><description>Meta is shutting down Facebook&apos;s face recognition system and deleting more than a billion face templates. An industry chapter closing, from someone who lived in it.</description><pubDate>Sat, 06 Nov 2021 00:00:00 GMT</pubDate><category>faces</category><category>privacy</category><category>policy</category></item><item><title>The metaverse needs faces first</title><link>https://vinsonli.com/posts/the-metaverse-needs-faces-first/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-metaverse-needs-faces-first/</guid><description>Facebook is now Meta and says it&apos;s building the metaverse. From someone who spent years on 3D faces: avatars that can&apos;t emote are just chat rooms with legs.</description><pubDate>Sat, 30 Oct 2021 00:00:00 GMT</pubDate><category>faces</category><category>vr</category><category>avatars</category><category>3d</category></item><item><title>Listening sessions are sentences</title><link>https://vinsonli.com/posts/listening-sessions-are-sentences/</link><guid isPermaLink="true">https://vinsonli.com/posts/listening-sessions-are-sentences/</guid><description>Treat each song as a token and each listening session as a sentence, and recommendation starts to look like language modeling. What that framing gets right, and what it misses.</description><pubDate>Mon, 27 Sep 2021 00:00:00 GMT</pubDate><category>recommendation</category><category>transformer</category><category>music</category><category>sequence-models</category></item><item><title>A dancer in a spandex suit, and why the humanoid form matters</title><link>https://vinsonli.com/posts/why-the-humanoid-form-matters/</link><guid isPermaLink="true">https://vinsonli.com/posts/why-the-humanoid-form-matters/</guid><description>Tesla&apos;s AI Day announced a humanoid robot and then showed a person dancing in a robot costume. The demo was a joke. The idea is right.</description><pubDate>Mon, 23 Aug 2021 00:00:00 GMT</pubDate><category>robotics</category><category>embodiment</category><category>tesla</category><category>humanoids</category></item><item><title>You can&apos;t skim a song</title><link>https://vinsonli.com/posts/you-cant-skim-a-song/</link><guid isPermaLink="true">https://vinsonli.com/posts/you-cant-skim-a-song/</guid><description>Why music recommendation is a different problem from video or news recommendation: sequences, repetition, passive listening, and the difference between a skip and a mood.</description><pubDate>Sun, 18 Jul 2021 00:00:00 GMT</pubDate><category>music</category><category>recommendation</category><category>product</category></item><item><title>Copilot finished my function</title><link>https://vinsonli.com/posts/copilot-finished-my-function/</link><guid isPermaLink="true">https://vinsonli.com/posts/copilot-finished-my-function/</guid><description>GitHub&apos;s Copilot preview suggests whole functions inside the editor. Two days of using it on side projects, and what changes when the IDE becomes a conversation.</description><pubDate>Tue, 06 Jul 2021 00:00:00 GMT</pubDate><category>coding</category><category>language</category><category>tools</category><category>openai</category></item><item><title>Joining YouTube Music: the recommendation is the product</title><link>https://vinsonli.com/posts/joining-youtube-music/</link><guid isPermaLink="true">https://vinsonli.com/posts/joining-youtube-music/</guid><description>I&apos;ve started as an engineering manager at YouTube Music. Why music, and why I think the next few years of music discovery will look very different.</description><pubDate>Wed, 26 May 2021 00:00:00 GMT</pubDate><category>career</category><category>music</category><category>recommendation</category><category>youtube</category></item><item><title>Closing two chapters</title><link>https://vinsonli.com/posts/closing-two-chapters/</link><guid isPermaLink="true">https://vinsonli.com/posts/closing-two-chapters/</guid><description>I&apos;m leaving TTT Studios after seven and a half years and Amanda AI after three. What I learned growing an engineering team from ten to fifty, and co-founding a company through a pandemic.</description><pubDate>Thu, 29 Apr 2021 00:00:00 GMT</pubDate><category>career</category><category>management</category><category>startups</category><category>amanda-ai</category></item><item><title>One architecture, any input</title><link>https://vinsonli.com/posts/one-architecture-any-input/</link><guid isPermaLink="true">https://vinsonli.com/posts/one-architecture-any-input/</guid><description>DeepMind&apos;s Perceiver handles images, audio, video and point clouds with the same network, by cross-attending to a small latent array. Modalities stop needing their own models.</description><pubDate>Sun, 14 Mar 2021 00:00:00 GMT</pubDate><category>architecture</category><category>multimodal</category><category>transformer</category><category>deepmind</category></item><item><title>Ten years between albums</title><link>https://vinsonli.com/posts/ten-years-between-albums/</link><guid isPermaLink="true">https://vinsonli.com/posts/ten-years-between-albums/</guid><description>万能青年旅店 released their second album after ten years. On listening to a band that sings about industrial decline in Hebei, and why art gets better once you can see the structure under it.</description><pubDate>Sun, 07 Feb 2021 00:00:00 GMT</pubDate><category>music</category><category>china</category><category>culture</category></item><item><title>Pictures and words in the same space</title><link>https://vinsonli.com/posts/pictures-and-words-in-the-same-space/</link><guid isPermaLink="true">https://vinsonli.com/posts/pictures-and-words-in-the-same-space/</guid><description>OpenAI&apos;s CLIP learns from 400 million image and caption pairs to put images and text in one embedding space. Zero-shot classification is the demo. Shared embeddings are the real story.</description><pubDate>Mon, 11 Jan 2021 00:00:00 GMT</pubDate><category>multimodal</category><category>representation</category><category>computer-vision</category><category>language</category></item><item><title>AlphaFold and learning physics from data</title><link>https://vinsonli.com/posts/alphafold-and-learning-physics-from-data/</link><guid isPermaLink="true">https://vinsonli.com/posts/alphafold-and-learning-physics-from-data/</guid><description>DeepMind&apos;s AlphaFold 2 predicts protein structures about as accurately as experiments do. A fifty-year-old physics problem, solved mostly by learning from examples.</description><pubDate>Thu, 03 Dec 2020 00:00:00 GMT</pubDate><category>science</category><category>deepmind</category><category>biology</category><category>physics</category></item><item><title>An image is worth 16x16 words</title><link>https://vinsonli.com/posts/an-image-is-worth-16x16-words/</link><guid isPermaLink="true">https://vinsonli.com/posts/an-image-is-worth-16x16-words/</guid><description>A paper under review at ICLR cuts images into patches and feeds them to a plain Transformer. With enough data, it beats convolutional networks. One architecture for everything is getting closer.</description><pubDate>Mon, 09 Nov 2020 00:00:00 GMT</pubDate><category>computer-vision</category><category>transformer</category><category>architecture</category><category>deep-learning</category></item><item><title>A whole scene stored inside a network</title><link>https://vinsonli.com/posts/a-whole-scene-stored-inside-a-network/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-whole-scene-stored-inside-a-network/</guid><description>NeRF represents a 3D scene as a small neural network you can render from any viewpoint. No mesh, no triangles. My old thesis problem, answered from a different direction.</description><pubDate>Thu, 08 Oct 2020 00:00:00 GMT</pubDate><category>3d</category><category>graphics</category><category>representation</category><category>deep-learning</category></item><item><title>Short video starts in India, which is exactly right</title><link>https://vinsonli.com/posts/short-video-starts-in-india/</link><guid isPermaLink="true">https://vinsonli.com/posts/short-video-starts-in-india/</guid><description>YouTube opened a beta of Shorts in India, where TikTok has been banned since June. Launch where the gap is, with the world&apos;s biggest music catalog behind you.</description><pubDate>Sat, 19 Sep 2020 00:00:00 GMT</pubDate><category>short-video</category><category>youtube</category><category>product</category><category>india</category></item><item><title>Reels is a clone. That&apos;s fine</title><link>https://vinsonli.com/posts/reels-is-a-clone-thats-fine/</link><guid isPermaLink="true">https://vinsonli.com/posts/reels-is-a-clone-thats-fine/</guid><description>Instagram launched Reels the day before the TikTok executive order. The format is a commodity now, and the question is whose recommender learns fastest.</description><pubDate>Tue, 11 Aug 2020 00:00:00 GMT</pubDate><category>short-video</category><category>recommendation</category><category>product</category><category>social</category></item><item><title>GPT-3 wrote a React component from a sentence</title><link>https://vinsonli.com/posts/gpt-3-and-the-react-component/</link><guid isPermaLink="true">https://vinsonli.com/posts/gpt-3-and-the-react-component/</guid><description>The demos flooding Twitter this week come from the same next-word objective as GPT-2, scaled a hundred times. Few-shot learning appeared on its own. Coding changes first.</description><pubDate>Tue, 21 Jul 2020 00:00:00 GMT</pubDate><category>language</category><category>gpt-3</category><category>coding</category><category>openai</category></item><item><title>IBM, Amazon and Microsoft step back from face recognition</title><link>https://vinsonli.com/posts/big-tech-steps-back-from-face-recognition/</link><guid isPermaLink="true">https://vinsonli.com/posts/big-tech-steps-back-from-face-recognition/</guid><description>In one week, three of the biggest companies in the field paused or ended police sales. Where I agree, and why a pause isn&apos;t a policy.</description><pubDate>Sat, 13 Jun 2020 00:00:00 GMT</pubDate><category>faces</category><category>policy</category><category>ethics</category><category>surveillance</category></item><item><title>What a face company does when nobody gathers</title><link>https://vinsonli.com/posts/what-a-face-company-does-when-nobody-gathers/</link><guid isPermaLink="true">https://vinsonli.com/posts/what-a-face-company-does-when-nobody-gathers/</guid><description>Three months into the pandemic, with no conferences on the calendar. What we&apos;re trying, what we decided not to build, and how it feels to run a startup whose market disappeared.</description><pubDate>Tue, 19 May 2020 00:00:00 GMT</pubDate><category>amanda-ai</category><category>startups</category><category>covid</category><category>product</category></item><item><title>Travis Scott played to 12 million people in a video game</title><link>https://vinsonli.com/posts/travis-scott-played-a-video-game/</link><guid isPermaLink="true">https://vinsonli.com/posts/travis-scott-played-a-video-game/</guid><description>Fortnite&apos;s Astronomical event drew 12 million concurrent players. With every venue closed, a game became the biggest concert stage in the world.</description><pubDate>Tue, 28 Apr 2020 00:00:00 GMT</pubDate><category>music</category><category>games</category><category>virtual-events</category><category>covid</category></item><item><title>Video is the new floor</title><link>https://vinsonli.com/posts/video-is-the-new-floor/</link><guid isPermaLink="true">https://vinsonli.com/posts/video-is-the-new-floor/</guid><description>Two weeks into working from home, every interaction is a video call. What that does to expectations for video, and what&apos;s missing from the tools.</description><pubDate>Fri, 27 Mar 2020 00:00:00 GMT</pubDate><category>video</category><category>covid</category><category>remote-work</category><category>product</category></item><item><title>MWC is cancelled. Our entire business is people in a room</title><link>https://vinsonli.com/posts/mwc-is-cancelled/</link><guid isPermaLink="true">https://vinsonli.com/posts/mwc-is-cancelled/</guid><description>The world&apos;s biggest mobile conference was cancelled over the new coronavirus. A founder&apos;s worries, a few weeks in, with friends in Wuhan under lockdown.</description><pubDate>Fri, 14 Feb 2020 00:00:00 GMT</pubDate><category>amanda-ai</category><category>startups</category><category>covid</category><category>events</category></item><item><title>Clearview scraped three billion faces. This is why we built on-device</title><link>https://vinsonli.com/posts/clearview-scraped-three-billion-faces/</link><guid isPermaLink="true">https://vinsonli.com/posts/clearview-scraped-three-billion-faces/</guid><description>The New York Times revealed a startup that matches any photo against billions of faces scraped from social media, and sells it to police. The worst version of my industry is now public.</description><pubDate>Tue, 21 Jan 2020 00:00:00 GMT</pubDate><category>faces</category><category>privacy</category><category>surveillance</category><category>amanda-ai</category></item><item><title>The decade cameras learned to see</title><link>https://vinsonli.com/posts/the-decade-cameras-learned-to-see/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-decade-cameras-learned-to-see/</guid><description>Looking back at 2010 to 2019 from the end of it: the moments that mattered to me, what I got right and wrong on this blog, and one bet for the 2020s.</description><pubDate>Mon, 30 Dec 2019 00:00:00 GMT</pubDate><category>retrospective</category><category>ai</category><category>predictions</category></item><item><title>MuZero learns the rules it isn&apos;t given</title><link>https://vinsonli.com/posts/muzero-learns-the-rules-it-isnt-given/</link><guid isPermaLink="true">https://vinsonli.com/posts/muzero-learns-the-rules-it-isnt-given/</guid><description>DeepMind&apos;s MuZero plays Go, chess, shogi and Atari at top level without being told the rules. It plans inside a model it learned, and the model only predicts what matters.</description><pubDate>Thu, 28 Nov 2019 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>world-models</category><category>deepmind</category><category>planning</category></item><item><title>200 seconds vs. 10,000 years</title><link>https://vinsonli.com/posts/200-seconds-vs-10000-years/</link><guid isPermaLink="true">https://vinsonli.com/posts/200-seconds-vs-10000-years/</guid><description>Google claims quantum supremacy with a 53-qubit chip. What the chip actually computed, why IBM disputes the number, and why there&apos;s still no workload.</description><pubDate>Sat, 26 Oct 2019 00:00:00 GMT</pubDate><category>quantum</category><category>hardware</category><category>physics</category><category>google</category></item><item><title>ZAO put your face in a movie in eight seconds</title><link>https://vinsonli.com/posts/zao-put-your-face-in-a-movie/</link><guid isPermaLink="true">https://vinsonli.com/posts/zao-put-your-face-in-a-movie/</guid><description>A Chinese face-swap app went viral over the weekend and hit a privacy wall within days. China just ran the consumer deepfake experiment at scale.</description><pubDate>Tue, 03 Sep 2019 00:00:00 GMT</pubDate><category>faces</category><category>deepfakes</category><category>china</category><category>privacy</category></item><item><title>Why our models train on forty different lobbies</title><link>https://vinsonli.com/posts/forty-different-lobbies/</link><guid isPermaLink="true">https://vinsonli.com/posts/forty-different-lobbies/</guid><description>A year of live events taught us that benchmark accuracy barely predicts real-world accuracy. What we changed in how we collect data and evaluate.</description><pubDate>Wed, 14 Aug 2019 00:00:00 GMT</pubDate><category>amanda-ai</category><category>machine-learning</category><category>data</category><category>operations</category></item><item><title>The embedding table is the model</title><link>https://vinsonli.com/posts/the-embedding-table-is-the-model/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-embedding-table-is-the-model/</guid><description>Facebook open-sourced DLRM, its deep learning recommendation model. It shows what big recommenders actually look like: mostly memory, with a small network on top.</description><pubDate>Sat, 20 Jul 2019 00:00:00 GMT</pubDate><category>recommendation</category><category>deep-learning</category><category>infrastructure</category></item><item><title>Datasets have a consent debt</title><link>https://vinsonli.com/posts/datasets-have-a-consent-debt/</link><guid isPermaLink="true">https://vinsonli.com/posts/datasets-have-a-consent-debt/</guid><description>Microsoft quietly took down MS-Celeb-1M, ten million photos of a hundred thousand people. A lot of the face recognition industry was built on data nobody agreed to give.</description><pubDate>Tue, 11 Jun 2019 00:00:00 GMT</pubDate><category>faces</category><category>privacy</category><category>datasets</category><category>ethics</category></item><item><title>San Francisco banned face recognition. Customers asked if we&apos;re next</title><link>https://vinsonli.com/posts/san-francisco-banned-face-recognition/</link><guid isPermaLink="true">https://vinsonli.com/posts/san-francisco-banned-face-recognition/</guid><description>The ban covers city agencies, not private companies like ours. It still changed a lot of sales conversations this week. Where I think the line should be.</description><pubDate>Fri, 17 May 2019 00:00:00 GMT</pubDate><category>faces</category><category>privacy</category><category>policy</category><category>amanda-ai</category></item><item><title>The black hole photo was a reconstruction problem</title><link>https://vinsonli.com/posts/the-black-hole-photo-was-a-reconstruction-problem/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-black-hole-photo-was-a-reconstruction-problem/</guid><description>The Event Horizon Telescope didn&apos;t take a picture in the ordinary sense. It filled in a mostly empty measurement with priors, and was careful about which ones.</description><pubDate>Sat, 13 Apr 2019 00:00:00 GMT</pubDate><category>imaging</category><category>science</category><category>computer-vision</category></item><item><title>The Bitter Lesson, read by someone who hand-built facial features</title><link>https://vinsonli.com/posts/the-bitter-lesson-read-by-a-hand-builder/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-bitter-lesson-read-by-a-hand-builder/</guid><description>Rich Sutton says 70 years of AI research show that general methods plus computation beat human knowledge every time. He&apos;s right, and it stings. My one objection is about data.</description><pubDate>Mon, 18 Mar 2019 00:00:00 GMT</pubDate><category>ai</category><category>research</category><category>essays</category></item><item><title>&quot;Too dangerous to release&quot;</title><link>https://vinsonli.com/posts/too-dangerous-to-release/</link><guid isPermaLink="true">https://vinsonli.com/posts/too-dangerous-to-release/</guid><description>OpenAI trained a much bigger language model and is holding back the full version. The unicorn story is impressive. What&apos;s actually dangerous is cheap, plausible text at scale.</description><pubDate>Tue, 19 Feb 2019 00:00:00 GMT</pubDate><category>language</category><category>ai-safety</category><category>openai</category><category>generative-models</category></item><item><title>AlphaStar and the fairness of fast hands</title><link>https://vinsonli.com/posts/alphastar-and-the-fairness-of-fast-hands/</link><guid isPermaLink="true">https://vinsonli.com/posts/alphastar-and-the-fairness-of-fast-hands/</guid><description>DeepMind&apos;s StarCraft agent beat two pros 10-0, then lost the one game where it had to move a camera like a person. What counts as intelligence when the body is different.</description><pubDate>Sun, 27 Jan 2019 00:00:00 GMT</pubDate><category>ai</category><category>games</category><category>deepmind</category><category>embodiment</category></item><item><title>None of these people exist</title><link>https://vinsonli.com/posts/none-of-these-people-exist/</link><guid isPermaLink="true">https://vinsonli.com/posts/none-of-these-people-exist/</guid><description>Nvidia&apos;s StyleGAN generates photographic faces with control over pose, identity and freckles. A face company&apos;s view of faces becoming free.</description><pubDate>Wed, 19 Dec 2018 00:00:00 GMT</pubDate><category>faces</category><category>generative-models</category><category>gans</category></item><item><title>The noisy TV problem</title><link>https://vinsonli.com/posts/the-noisy-tv-problem/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-noisy-tv-problem/</guid><description>A curious agent gets hypnotized by random static. OpenAI&apos;s Random Network Distillation fixes it and beats humans at Montezuma&apos;s Revenge. What curiosity should actually reward.</description><pubDate>Sat, 17 Nov 2018 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>curiosity</category><category>exploration</category></item><item><title>BERT reads both directions at once</title><link>https://vinsonli.com/posts/bert-reads-both-directions/</link><guid isPermaLink="true">https://vinsonli.com/posts/bert-reads-both-directions/</guid><description>Google&apos;s BERT beat almost every language benchmark by predicting hidden words using context on both sides. It&apos;s a representation model, which is a different thing from a text generator.</description><pubDate>Mon, 29 Oct 2018 00:00:00 GMT</pubDate><category>language</category><category>transformer</category><category>pretraining</category><category>representation</category></item><item><title>Thousands of faces in a hotel lobby</title><link>https://vinsonli.com/posts/thousands-of-faces-in-a-hotel-lobby/</link><guid isPermaLink="true">https://vinsonli.com/posts/thousands-of-faces-in-a-hotel-lobby/</guid><description>Our first live conference check-ins. Wait times dropped a lot. The model was the part that worked from day one. Lighting, queues and printers were not.</description><pubDate>Tue, 18 Sep 2018 00:00:00 GMT</pubDate><category>amanda-ai</category><category>startups</category><category>faces</category><category>operations</category></item><item><title>Musical.ly is now TikTok, as predicted</title><link>https://vinsonli.com/posts/musically-is-now-tiktok/</link><guid isPermaLink="true">https://vinsonli.com/posts/musically-is-now-tiktok/</guid><description>ByteDance merged Musical.ly into TikTok this week. A short follow-up to what I wrote in November, and what the feed tells you about the product.</description><pubDate>Wed, 08 Aug 2018 00:00:00 GMT</pubDate><category>recommendation</category><category>short-video</category><category>china</category><category>product</category></item><item><title>A robot hand with a hundred years of practice</title><link>https://vinsonli.com/posts/a-robot-hand-with-a-hundred-years-of-practice/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-robot-hand-with-a-hundred-years-of-practice/</guid><description>OpenAI&apos;s Dactyl learned to rotate a block with a human-like robot hand, entirely in simulation. Why hands, and why randomizing the simulator is the clever part.</description><pubDate>Tue, 31 Jul 2018 00:00:00 GMT</pubDate><category>robotics</category><category>hands</category><category>simulation</category><category>reinforcement-learning</category></item><item><title>Read everything first, specialize later</title><link>https://vinsonli.com/posts/read-everything-first-specialize-later/</link><guid isPermaLink="true">https://vinsonli.com/posts/read-everything-first-specialize-later/</guid><description>OpenAI trained a Transformer to predict the next word on thousands of books, then fine-tuned it on small tasks. Language is getting its ImageNet moment.</description><pubDate>Wed, 20 Jun 2018 00:00:00 GMT</pubDate><category>language</category><category>transformer</category><category>pretraining</category><category>deep-learning</category></item><item><title>The Uber car saw her six seconds early</title><link>https://vinsonli.com/posts/the-uber-car-saw-her-six-seconds-early/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-uber-car-saw-her-six-seconds-early/</guid><description>The NTSB&apos;s preliminary report on the Tempe crash shows the sensors worked. Classification flickered, braking was disabled, and the human was supposed to be the backup.</description><pubDate>Sun, 27 May 2018 00:00:00 GMT</pubDate><category>self-driving</category><category>safety</category><category>systems</category></item><item><title>Faces help you hear</title><link>https://vinsonli.com/posts/faces-help-you-hear/</link><guid isPermaLink="true">https://vinsonli.com/posts/faces-help-you-hear/</guid><description>Google&apos;s Looking to Listen separates one voice from a crowd by watching the speaker&apos;s face. Modalities work better when they explain each other.</description><pubDate>Sat, 21 Apr 2018 00:00:00 GMT</pubDate><category>faces</category><category>audio</category><category>multimodal</category><category>deep-learning</category></item><item><title>An agent that dreams its own racetrack</title><link>https://vinsonli.com/posts/an-agent-that-dreams-its-own-racetrack/</link><guid isPermaLink="true">https://vinsonli.com/posts/an-agent-that-dreams-its-own-racetrack/</guid><description>Ha and Schmidhuber&apos;s World Models compresses what an agent sees, learns to predict what happens next, and trains a tiny controller inside its own dream. The most important paper I&apos;ve read this year.</description><pubDate>Sat, 31 Mar 2018 00:00:00 GMT</pubDate><category>world-models</category><category>reinforcement-learning</category><category>generative-models</category></item><item><title>Starting a face recognition company, with eyes open</title><link>https://vinsonli.com/posts/starting-a-face-recognition-company/</link><guid isPermaLink="true">https://vinsonli.com/posts/starting-a-face-recognition-company/</guid><description>I&apos;ve co-founded a company that does face recognition on device. Why on device, why conference check-in first, and the lines we won&apos;t cross.</description><pubDate>Mon, 26 Feb 2018 00:00:00 GMT</pubDate><category>startups</category><category>faces</category><category>privacy</category><category>amanda-ai</category></item><item><title>Can a network have taste?</title><link>https://vinsonli.com/posts/can-a-network-have-taste/</link><guid isPermaLink="true">https://vinsonli.com/posts/can-a-network-have-taste/</guid><description>Google&apos;s NIMA predicts how people would rate a photo, as a distribution, not a single score. Modeling disagreement turns out to be the useful part.</description><pubDate>Mon, 22 Jan 2018 00:00:00 GMT</pubDate><category>aesthetics</category><category>deep-learning</category><category>photography</category><category>taste</category></item><item><title>Face swaps on Reddit: I built the harmless version years ago</title><link>https://vinsonli.com/posts/face-swaps-on-reddit/</link><guid isPermaLink="true">https://vinsonli.com/posts/face-swaps-on-reddit/</guid><description>Someone on Reddit is putting celebrities&apos; faces into porn with a home GPU and open-source tools. How it works, and why consent has to be designed in now.</description><pubDate>Sat, 16 Dec 2017 00:00:00 GMT</pubDate><category>faces</category><category>deepfakes</category><category>ethics</category><category>generative-models</category></item><item><title>Watch the recommendation engine, not the lip-sync</title><link>https://vinsonli.com/posts/watch-the-recommendation-engine/</link><guid isPermaLink="true">https://vinsonli.com/posts/watch-the-recommendation-engine/</guid><description>ByteDance is buying Musical.ly. It looks like a Chinese news company buying an American teen app. It&apos;s really a recommendation engine buying a global audience.</description><pubDate>Tue, 14 Nov 2017 00:00:00 GMT</pubDate><category>recommendation</category><category>short-video</category><category>china</category><category>product</category></item><item><title>AlphaGo Zero threw away human games and got better</title><link>https://vinsonli.com/posts/alphago-zero-threw-away-human-games/</link><guid isPermaLink="true">https://vinsonli.com/posts/alphago-zero-threw-away-human-games/</guid><description>Starting from random play, with no human data, it beat the version that beat Lee Sedol 100 games to 0 after three days. Human knowledge was a ceiling.</description><pubDate>Sat, 21 Oct 2017 00:00:00 GMT</pubDate><category>ai</category><category>reinforcement-learning</category><category>deepmind</category><category>go</category></item><item><title>Face ID: the 3D face is finally consumer hardware</title><link>https://vinsonli.com/posts/face-id-the-3d-face-is-consumer-hardware/</link><guid isPermaLink="true">https://vinsonli.com/posts/face-id-the-3d-face-is-consumer-hardware/</guid><description>The iPhone X projects 30,000 infrared dots onto your face to unlock the phone. Four years ago I was faking this in software from a single selfie.</description><pubDate>Wed, 13 Sep 2017 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>hardware</category><category>apple</category><category>security</category></item><item><title>Nobody taught it to walk</title><link>https://vinsonli.com/posts/nobody-taught-it-to-walk/</link><guid isPermaLink="true">https://vinsonli.com/posts/nobody-taught-it-to-walk/</guid><description>DeepMind&apos;s locomotion agents learned to run, jump and duck from a reward for moving forward and a varied course. The flailing arms are the point.</description><pubDate>Sat, 12 Aug 2017 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>robotics</category><category>simulation</category><category>embodiment</category></item><item><title>No recurrence, no convolution</title><link>https://vinsonli.com/posts/no-recurrence-no-convolution/</link><guid isPermaLink="true">https://vinsonli.com/posts/no-recurrence-no-convolution/</guid><description>A Google paper throws out the RNN and translates with attention alone. What attention actually computes, and why I think it goes beyond translation.</description><pubDate>Thu, 06 Jul 2017 00:00:00 GMT</pubDate><category>deep-learning</category><category>language</category><category>architecture</category><category>transformer</category></item><item><title>Curiosity is a prediction error</title><link>https://vinsonli.com/posts/curiosity-is-a-prediction-error/</link><guid isPermaLink="true">https://vinsonli.com/posts/curiosity-is-a-prediction-error/</guid><description>A Berkeley paper gives agents an internal reward for being surprised by their own predictions. It learns Mario with no score at all, and it answers a question I asked two years ago.</description><pubDate>Sat, 17 Jun 2017 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>curiosity</category><category>learning</category></item><item><title>Ke Jie in Wuzhen</title><link>https://vinsonli.com/posts/ke-jie-in-wuzhen/</link><guid isPermaLink="true">https://vinsonli.com/posts/ke-jie-in-wuzhen/</guid><description>The world&apos;s top Go player lost 3-0 to AlphaGo in China this week, and Chinese viewers mostly couldn&apos;t watch it live. Why this match will matter more in China than the Lee Sedol one did.</description><pubDate>Sun, 28 May 2017 00:00:00 GMT</pubDate><category>ai</category><category>go</category><category>china</category><category>deepmind</category></item><item><title>Neuralink and the bandwidth argument</title><link>https://vinsonli.com/posts/neuralink-and-the-bandwidth-argument/</link><guid isPermaLink="true">https://vinsonli.com/posts/neuralink-and-the-bandwidth-argument/</guid><description>Elon Musk&apos;s new company wants a high-bandwidth link between brains and computers. The argument is about output speed. I think it&apos;s partly right, and too optimistic on timing.</description><pubDate>Sun, 23 Apr 2017 00:00:00 GMT</pubDate><category>bci</category><category>hardware</category><category>neuroscience</category></item><item><title>Breath of the Wild is a physics engine you can think in</title><link>https://vinsonli.com/posts/breath-of-the-wild-is-a-physics-engine/</link><guid isPermaLink="true">https://vinsonli.com/posts/breath-of-the-wild-is-a-physics-engine/</guid><description>The new Zelda replaced scripted puzzles with a world that follows consistent rules. It&apos;s the best example I&apos;ve seen of what a training environment for intelligence could look like.</description><pubDate>Sat, 18 Mar 2017 00:00:00 GMT</pubDate><category>games</category><category>physics</category><category>simulation</category><category>learning</category></item><item><title>The camera is the new keyboard</title><link>https://vinsonli.com/posts/the-camera-is-the-new-keyboard/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-camera-is-the-new-keyboard/</guid><description>Snap filed to go public and calls itself a camera company. The phrase sounds like marketing, but it describes a real shift in how people put things into computers.</description><pubDate>Tue, 07 Feb 2017 00:00:00 GMT</pubDate><category>camera</category><category>social</category><category>product</category><category>faces</category></item><item><title>Poker fell</title><link>https://vinsonli.com/posts/poker-fell/</link><guid isPermaLink="true">https://vinsonli.com/posts/poker-fell/</guid><description>CMU&apos;s Libratus beat four top professionals at heads-up no-limit hold&apos;em over twenty days. Hidden information was supposed to protect humans a while longer.</description><pubDate>Tue, 31 Jan 2017 00:00:00 GMT</pubDate><category>ai</category><category>games</category><category>game-theory</category></item><item><title>Every lab is building a playground</title><link>https://vinsonli.com/posts/every-lab-is-building-a-playground/</link><guid isPermaLink="true">https://vinsonli.com/posts/every-lab-is-building-a-playground/</guid><description>DeepMind open-sourced its 3D lab and OpenAI released Universe in the same week. Environments are becoming the new datasets.</description><pubDate>Fri, 09 Dec 2016 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>simulation</category><category>research</category></item><item><title>Google Translate got better overnight, and Chinese speakers noticed first</title><link>https://vinsonli.com/posts/google-translate-got-better-overnight/</link><guid isPermaLink="true">https://vinsonli.com/posts/google-translate-got-better-overnight/</guid><description>Neural machine translation replaced phrase tables, starting with Chinese to English. Notes from someone who reads both, and why the zero-shot result is the bigger story.</description><pubDate>Sat, 19 Nov 2016 00:00:00 GMT</pubDate><category>translation</category><category>language</category><category>deep-learning</category><category>chinese</category></item><item><title>Vine is dead. Short video isn&apos;t</title><link>https://vinsonli.com/posts/vine-is-dead-short-video-isnt/</link><guid isPermaLink="true">https://vinsonli.com/posts/vine-is-dead-short-video-isnt/</guid><description>Twitter is shutting down Vine. It didn&apos;t lose because six-second videos were a bad idea.</description><pubDate>Sat, 29 Oct 2016 00:00:00 GMT</pubDate><category>video</category><category>social</category><category>product</category><category>creators</category></item><item><title>16,000 samples a second</title><link>https://vinsonli.com/posts/16000-samples-a-second/</link><guid isPermaLink="true">https://vinsonli.com/posts/16000-samples-a-second/</guid><description>DeepMind&apos;s WaveNet generates raw audio one sample at a time, and its speech sounds far more human than anything before. How, and why it&apos;s so slow.</description><pubDate>Wed, 14 Sep 2016 00:00:00 GMT</pubDate><category>audio</category><category>generative-models</category><category>deepmind</category><category>speech</category></item><item><title>The long tail of driving</title><link>https://vinsonli.com/posts/the-long-tail-of-driving/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-long-tail-of-driving/</guid><description>Uber will put self-driving cars on Pittsburgh streets next month, six weeks after the first Autopilot fatality became public. The hard part of driving is the rare case.</description><pubDate>Sun, 21 Aug 2016 00:00:00 GMT</pubDate><category>self-driving</category><category>robotics</category><category>safety</category></item><item><title>Pokémon Go is AR without the AR</title><link>https://vinsonli.com/posts/pokemon-go-is-ar-without-the-ar/</link><guid isPermaLink="true">https://vinsonli.com/posts/pokemon-go-is-ar-without-the-ar/</guid><description>The biggest augmented reality hit so far barely uses augmented reality. It&apos;s a map, built on years of data from another game.</description><pubDate>Wed, 13 Jul 2016 00:00:00 GMT</pubDate><category>ar</category><category>games</category><category>product</category><category>maps</category></item><item><title>Still caring about facial conformation</title><link>https://vinsonli.com/posts/still-caring-about-facial-conformation/</link><guid isPermaLink="true">https://vinsonli.com/posts/still-caring-about-facial-conformation/</guid><description>My last paper from grad school is out. Why I still think geometric priors for faces matter when every result now comes from deep learning.</description><pubDate>Thu, 30 Jun 2016 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>research</category></item><item><title>Google built its own chip for neural nets</title><link>https://vinsonli.com/posts/google-built-its-own-chip/</link><guid isPermaLink="true">https://vinsonli.com/posts/google-built-its-own-chip/</guid><description>The Tensor Processing Unit has been running in Google&apos;s data centers for a year. When a company designs custom silicon for a workload, the workload is here to stay.</description><pubDate>Sun, 22 May 2016 00:00:00 GMT</pubDate><category>hardware</category><category>machine-learning</category><category>google</category><category>infrastructure</category></item><item><title>Chatbots are this year&apos;s apps. Tay says slow down</title><link>https://vinsonli.com/posts/chatbots-are-this-years-apps/</link><guid isPermaLink="true">https://vinsonli.com/posts/chatbots-are-this-years-apps/</guid><description>Facebook opened Messenger to bots at F8, two weeks after Microsoft&apos;s Tay learned racism from Twitter in a day. Why I&apos;m telling clients to wait.</description><pubDate>Thu, 14 Apr 2016 00:00:00 GMT</pubDate><category>chatbots</category><category>language</category><category>product</category><category>business</category></item><item><title>Move 37</title><link>https://vinsonli.com/posts/move-37/</link><guid isPermaLink="true">https://vinsonli.com/posts/move-37/</guid><description>AlphaGo beat Lee Sedol four games to one. The move everyone is talking about, how the system found it, and why self-play is the part that matters.</description><pubDate>Wed, 16 Mar 2016 00:00:00 GMT</pubDate><category>ai</category><category>go</category><category>reinforcement-learning</category><category>deepmind</category></item><item><title>The best robotics video ever made involves a hockey stick</title><link>https://vinsonli.com/posts/the-best-robotics-video-involves-a-hockey-stick/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-best-robotics-video-involves-a-hockey-stick/</guid><description>Boston Dynamics&apos; new Atlas gets shoved, loses its box and gets knocked flat, and keeps going. Why recovering is harder than walking.</description><pubDate>Fri, 26 Feb 2016 00:00:00 GMT</pubDate><category>robotics</category><category>control</category><category>physics</category></item><item><title>Re-reading Society of Mind as an engineering manager</title><link>https://vinsonli.com/posts/society-of-mind-as-an-engineering-manager/</link><guid isPermaLink="true">https://vinsonli.com/posts/society-of-mind-as-an-engineering-manager/</guid><description>Marvin Minsky died this week, the same week DeepMind announced a Go program that beat a professional. Two very different ideas of how minds get built.</description><pubDate>Fri, 29 Jan 2016 00:00:00 GMT</pubDate><category>ai</category><category>books</category><category>management</category><category>go</category></item><item><title>152 layers, and the trick is learning nothing</title><link>https://vinsonli.com/posts/152-layers-and-the-trick-is-learning-nothing/</link><guid isPermaLink="true">https://vinsonli.com/posts/152-layers-and-the-trick-is-learning-nothing/</guid><description>Microsoft Research&apos;s residual networks won ImageNet with a network eight times deeper than last year&apos;s. The idea behind it is almost too simple.</description><pubDate>Tue, 15 Dec 2015 00:00:00 GMT</pubDate><category>deep-learning</category><category>computer-vision</category><category>architecture</category></item><item><title>TensorFlow is open source. What our studio will do with it</title><link>https://vinsonli.com/posts/tensorflow-is-open-source/</link><guid isPermaLink="true">https://vinsonli.com/posts/tensorflow-is-open-source/</guid><description>Google released its internal machine learning library. What changes for a thirty-person app studio that builds software for other companies.</description><pubDate>Thu, 12 Nov 2015 00:00:00 GMT</pubDate><category>machine-learning</category><category>tools</category><category>engineering</category><category>business</category></item><item><title>AC/DC at sixty</title><link>https://vinsonli.com/posts/ac-dc-at-sixty/</link><guid isPermaLink="true">https://vinsonli.com/posts/ac-dc-at-sixty/</guid><description>Angus Young is sixty and still runs across the stage in a school uniform. Some thoughts after the show.</description><pubDate>Sat, 17 Oct 2015 00:00:00 GMT</pubDate><category>music</category><category>life</category></item><item><title>The screen finally feels how hard you press</title><link>https://vinsonli.com/posts/the-screen-feels-how-hard-you-press/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-screen-feels-how-hard-you-press/</guid><description>3D Touch on the iPhone 6s, one year after I guessed pressure-sensitive phones were a couple of years away. How it measures force, and what it&apos;s for.</description><pubDate>Sat, 19 Sep 2015 00:00:00 GMT</pubDate><category>hardware</category><category>interfaces</category><category>haptics</category><category>apple</category></item><item><title>Van Gogh is a Gram matrix</title><link>https://vinsonli.com/posts/van-gogh-is-a-gram-matrix/</link><guid isPermaLink="true">https://vinsonli.com/posts/van-gogh-is-a-gram-matrix/</guid><description>A new paper separates the content of an image from its style using a network trained for classification. How it works, and what it suggests about taste.</description><pubDate>Mon, 31 Aug 2015 00:00:00 GMT</pubDate><category>deep-learning</category><category>art</category><category>style</category><category>generative-models</category></item><item><title>Apple bets on human DJs. I&apos;d bet on the skip button</title><link>https://vinsonli.com/posts/human-djs-vs-the-skip-button/</link><guid isPermaLink="true">https://vinsonli.com/posts/human-djs-vs-the-skip-button/</guid><description>Apple Music launched with Beats 1 and human curators. Three weeks later, Spotify ships Discover Weekly. Two theories of music discovery.</description><pubDate>Fri, 24 Jul 2015 00:00:00 GMT</pubDate><category>music</category><category>recommendation</category><category>product</category></item><item><title>DeepDream sees dogs everywhere</title><link>https://vinsonli.com/posts/deepdream-sees-dogs-everywhere/</link><guid isPermaLink="true">https://vinsonli.com/posts/deepdream-sees-dogs-everywhere/</guid><description>Running a network in reverse to see what it learned. Why everything turns into dogs, and what that says about training data.</description><pubDate>Fri, 26 Jun 2015 00:00:00 GMT</pubDate><category>deep-learning</category><category>computer-vision</category><category>visualization</category></item><item><title>Google Photos gives away what I spent two years building</title><link>https://vinsonli.com/posts/google-photos-gives-away-what-i-built/</link><guid isPermaLink="true">https://vinsonli.com/posts/google-photos-gives-away-what-i-built/</guid><description>Unlimited storage and automatic face grouping, free. What happens to small companies when a platform ships their feature as a default.</description><pubDate>Sun, 31 May 2015 00:00:00 GMT</pubDate><category>product</category><category>faces</category><category>platforms</category><category>google</category></item><item><title>Two years of 3D faces on a phone: what I got wrong</title><link>https://vinsonli.com/posts/two-years-of-3d-faces-what-i-got-wrong/</link><guid isPermaLink="true">https://vinsonli.com/posts/two-years-of-3d-faces-what-i-got-wrong/</guid><description>Leaving the selfie-to-3D-face app I helped build out of my thesis. The technology mostly worked. The rest is the interesting part.</description><pubDate>Thu, 30 Apr 2015 00:00:00 GMT</pubDate><category>startups</category><category>faces</category><category>product</category><category>career</category></item><item><title>The dress is a world-model bug</title><link>https://vinsonli.com/posts/the-dress-is-a-world-model-bug/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-dress-is-a-world-model-bug/</guid><description>Half the internet sees white and gold, half sees blue and black. The disagreement is about lighting, and it says something about how vision works.</description><pubDate>Tue, 03 Mar 2015 00:00:00 GMT</pubDate><category>perception</category><category>vision</category><category>color</category></item><item><title>49 Atari games from pixels, and zero points in Montezuma&apos;s Revenge</title><link>https://vinsonli.com/posts/zero-points-in-montezumas-revenge/</link><guid isPermaLink="true">https://vinsonli.com/posts/zero-points-in-montezumas-revenge/</guid><description>DeepMind&apos;s DQN paper is in Nature. What the network actually learns, and the game where it learns nothing at all.</description><pubDate>Fri, 27 Feb 2015 00:00:00 GMT</pubDate><category>reinforcement-learning</category><category>deep-learning</category><category>deepmind</category></item><item><title>HoloLens maps your room before it draws on it</title><link>https://vinsonli.com/posts/hololens-maps-your-room-first/</link><guid isPermaLink="true">https://vinsonli.com/posts/hololens-maps-your-room-first/</guid><description>Microsoft&apos;s headset is mostly a sensing problem with a display attached. Why AR needs a model of the space before it can put anything in it.</description><pubDate>Tue, 27 Jan 2015 00:00:00 GMT</pubDate><category>ar</category><category>hardware</category><category>computer-vision</category></item><item><title>Two networks arguing</title><link>https://vinsonli.com/posts/two-networks-arguing/</link><guid isPermaLink="true">https://vinsonli.com/posts/two-networks-arguing/</guid><description>The GAN paper from NIPS this year: a generator, a discriminator, and a learned idea of what counts as real.</description><pubDate>Thu, 18 Dec 2014 00:00:00 GMT</pubDate><category>deep-learning</category><category>generative-models</category><category>faces</category></item><item><title>A computer captioned a photo. It didn&apos;t see the photo</title><link>https://vinsonli.com/posts/a-computer-captioned-a-photo/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-computer-captioned-a-photo/</guid><description>Google and Stanford both have networks that write sentences about images. How they work, and what they&apos;re actually learning.</description><pubDate>Sat, 22 Nov 2014 00:00:00 GMT</pubDate><category>computer-vision</category><category>deep-learning</category><category>language</category></item><item><title>What a light field actually is</title><link>https://vinsonli.com/posts/what-a-light-field-actually-is/</link><guid isPermaLink="true">https://vinsonli.com/posts/what-a-light-field-actually-is/</guid><description>Magic Leap raised $542 million without showing anyone a product. A look at the optics they&apos;re promising, and why it&apos;s so hard to wear on your head.</description><pubDate>Sat, 25 Oct 2014 00:00:00 GMT</pubDate><category>ar</category><category>optics</category><category>hardware</category></item><item><title>The watch taps your wrist, and touch becomes a channel</title><link>https://vinsonli.com/posts/the-watch-taps-your-wrist/</link><guid isPermaLink="true">https://vinsonli.com/posts/the-watch-taps-your-wrist/</guid><description>The part of the Apple Watch announcement I keep thinking about is a small mass on a spring, and what it says about how little computers use touch.</description><pubDate>Tue, 16 Sep 2014 00:00:00 GMT</pubDate><category>haptics</category><category>hardware</category><category>interfaces</category><category>sensing</category></item><item><title>Radial basis functions, for people who don&apos;t care about radial basis functions</title><link>https://vinsonli.com/posts/radial-basis-functions-for-people-who-dont-care/</link><guid isPermaLink="true">https://vinsonli.com/posts/radial-basis-functions-for-people-who-dont-care/</guid><description>I presented a facial animation paper this week. What it&apos;s about, explained with a rubber sheet, some pins and a steel plate.</description><pubDate>Tue, 26 Aug 2014 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>math</category><category>research</category></item><item><title>Why my 3D faces look dead: the physics of skin</title><link>https://vinsonli.com/posts/why-my-3d-faces-look-dead/</link><guid isPermaLink="true">https://vinsonli.com/posts/why-my-3d-faces-look-dead/</guid><description>Our app gets the shape of your face right and still looks like a wax museum. Most of the problem turned out to be motion.</description><pubDate>Sat, 19 Jul 2014 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>animation</category><category>physics</category><category>product</category></item><item><title>A video is not a stack of photos</title><link>https://vinsonli.com/posts/a-video-is-not-a-stack-of-photos/</link><guid isPermaLink="true">https://vinsonli.com/posts/a-video-is-not-a-stack-of-photos/</guid><description>Two new papers on video classification, and why the network that only sees motion beat the one that sees the frames.</description><pubDate>Sun, 22 Jun 2014 00:00:00 GMT</pubDate><category>video</category><category>computer-vision</category><category>deep-learning</category><category>optical-flow</category></item><item><title>Graduated. My thesis in one paragraph, and how deep learning will eat it</title><link>https://vinsonli.com/posts/graduated-thesis-in-one-paragraph/</link><guid isPermaLink="true">https://vinsonli.com/posts/graduated-thesis-in-one-paragraph/</guid><description>What my master&apos;s thesis on 3D face meshes does, which part of it I think neural networks will replace soon, and which part I think they won&apos;t.</description><pubDate>Thu, 29 May 2014 00:00:00 GMT</pubDate><category>faces</category><category>3d</category><category>deep-learning</category><category>career</category></item></channel></rss>