Physical AI's Next Bottleneck Is Finding the Right Video

Rafael Levi, Bright Data16:32 · Sept 2026 · 4,127 views
Thumbnail for Physical AI's Next Bottleneck Is Finding the Right Video Watch on YouTube
TL;DR
  1. 1

    Robotics has far less training video than language and image models have training data, while staged recordings give robots unnatural examples.

  2. 2

    The public web contains useful footage of people handling objects and dealing with real physics, but most downloaded video is irrelevant to a specific action.

  3. 3

    Bright Data indexes videos by the actions inside them and returns trimmed clips, timestamps, match scores, and frame counts through an API.

Summary

Rafael Levi argues that robotics has reached a point where data collection is a larger problem than model design. Language models can train on trillions of words, while robotics has only about a million videos of robots doing things. Recordings made by people who are instructed to open doors or sit in chairs are biased because people move differently when they are performing for a camera. Simulation and teleoperation also have limits in physics, cost, and scale. Levi points to the public video web as a much larger source of natural actions, object handling, motion, gravity, and cause and effect. The problem is finding the useful parts. He cites NVIDIA discarding about 96% of downloaded video for Cosmos and Stable Video Diffusion discarding 74%. Bright Data's proposed solution is to search indexed video by actions, then collect only matching clips. The demo covers dishwashing, folding clothes, makeup use, brand discovery, self-driving footage, and physics.

Key ideas
01:00

Robotics has a data shortage even though its models are improving

Levi says the hard part of building systems that see, understand, and act is now the data provided to the AI. Language models have trillions of words and image generation has billions of labeled images, but robotics has only about a million videos of robots doing things. He places this shortage in the context of recent progress, from Google's robot learning work in 2022 to systems controlling multiple robots in 2023 and open-source models in 2024.

03:15

Instructed recordings teach robots unnatural movements

Companies pay people to record actions such as opening a door or sitting in a chair. Levi says people do not perform these actions naturally when they have been told to record them. Their movement changes because they are acting for someone and may also behave differently in front of a camera. He calls this biased data and argues that it will not give a robot the same results as footage of an intuitive, everyday action.

04:35

Simulation and teleoperation do not provide enough natural training data

Levi says simulation is cheap, but the physics in video games is not good enough for robot training. Teleoperation can produce recordings, but it is difficult to scale the number of hours and people required. A person might control a robot for eight hours a day, yet that still does not create enough data over a year. Existing pre-built datasets are also small, with about a million robot videos.

05:14

Public video contains actions and physical events robots need to learn

Levi points to the web as a large source of real-world footage. Videos show people handling objects, along with gravity, motion, and cause and effect. He cites a Meta model trained on about a million hours of public video that needed only 62 hours of real robot data to control a robot. In his account, the public footage supplied broad visual experience before the smaller robot dataset was used.

06:54

Frame-to-frame motion lets models learn from ordinary videos

A video does not need to show a robot for every useful action to be extracted. Levi explains that an AI can compare consecutive frames, measure the changing movement, and continue through the video to estimate angles, distance, and other information. The raw video still needs processing, possibly for motion, distance, and sensor-related information, but the useful material already exists in online videos.

07:58

Most downloaded video can be irrelevant to the training action

Levi says NVIDIA throws out about 96% of the video downloaded when training Cosmos, so only about 4% of a million hours may be useful in that example. He also says Stable Video Diffusion throws out 74% of the videos it downloads. The discarded material consumes compute, bandwidth, and storage before anyone can use it.

08:33

Bright Data searches video by actions before collecting clips

Bright Data's proposed workflow is 'search first, collect second.' The platform indexes more than a billion videos and searches them by actions rather than only by titles or keywords. A query such as a person washing dishes or folding a T-shirt returns trimmed snippets that are prepared for further processing. Users can add detailed descriptions to narrow the results.

11:38

The API returns evidence for selecting clips across several use cases

Levi says the system is available through an API. It returns a video snippet, the original URL, a timestamp, a match score, and a frame count. The demo searches for dishwashing and clothes folding, but he also describes finding makeup videos where a particular brand appears, locating driving situations in dashcam footage, and searching for other physical events such as an apple falling from a tree.

"You do not need to download million of hours of videos. You can actually get exactly the specific actions that you are interested in."13:04
Who should watch
  • You are building a robotics or world-model dataset and are spending time downloading broad collections of online video before filtering them.
  • Your robot training data comes from staged demonstrations, simulation, or teleoperation, and you want examples of natural human actions and real-world physics.
  • You work on self-driving, brand discovery, or video search and need to find events inside videos rather than relying on titles.