From VLM/VLAs to Embodied Agents

Armen Aghajanyan, Perceptron AI20:42 · Sept 2026 · 3,145 views
Thumbnail for From VLM/VLAs to Embodied Agents Watch on YouTube
TL;DR
  1. 1

    Perceptron wants one embodied foundation model that can perceive, reason, and control devices in the physical world.

  2. 2

    A perceptive objective can teach a model which visual details matter instead of applying equal weight to every pixel.

  3. 3

    Joint training on perception, reasoning, and control can reduce the amount of expensive teleoperation data needed for robotics.

Summary

Armen Aghajanyan argues that VLMs, embodied reasoning models, VLAs, and world models should be combined into one embodied foundation model. The model should take in different modalities, understand them, and produce outputs ranging from text and grounding points to control actions. He focuses on two problems. Video contains huge numbers of tokens but very little useful ground truth, while always-on cameras create excessive context. Perceptron's approach learns which percepts and tokens matter for a task. Its data-sparse mixture of experts lets a router skip tokens or allocate more computation to relevant regions. A model trained on a petabyte of text, images, video, and trajectories can treat detection as an agentic process, including tiling an image and changing contrast. A reported scaling result suggests that more video pretraining can substitute for teleoperation data. The talk ends with a single model reading book titles and emitting control tokens to sort them.

Key ideas
00:13

Perceptron wants perception, reasoning, and control in one model

Armen Aghajanyan describes Perceptron's goal as building physical AI foundations that can perceive, understand, and interact with the physical world in real time. The same system should work wherever there is a device, instrument, robot, camera, or sensor. He places VLMs, embodied reasoning models, VLAs, world models, and semantic world models on a progression of input and output capabilities. An embodied foundation model combines standard perception, embodied reasoning, and control within one model, with multiple modalities on the input side.

03:40

Video pretraining wastes most of its visual signal

A one-hour video can produce around one million visual tokens, but the available supervision may come only from transcripts, synthetic frame labels, or questions. Aghajanyan says this means the loss is calculated on roughly two percent of the incoming tokens. Predicting every pixel gives a dense signal, yet it assigns similar importance to background pixels and to details such as a gripper tip, contact points, failures, or physics. Perceptron's proposed direction is a natural perceptive objective that learns which percepts will matter later, rather than hardcoding a fixed percept such as a robot gripper.

05:41

Sparse mixture-of-experts routing lets the model choose visual tokens

Always-on cameras and continuously operating robots create long contexts filled with video tokens. Aghajanyan says text is relatively dense while video is relatively sparse, so the modalities should not be treated identically. Patch-wise averaging can provide compression of up to ten times, but he calls it a limited workaround. Perceptron's data-sparse mixture-of-experts model uses a router at each layer to predict which tokens to read and which to skip. Visualizations show it focusing on a graph in a figure, spreading attention for a general question, and allocating more tokens to fruit when asked to segment fruit.

09:50

Detection can be an agentic process rather than one fixed prediction

Perceptron's model treats a classical computer vision task such as detection as a sequence of decisions. It can write code, tile an image, zoom into regions, change contrast, and propose boxes before finding the target. Aghajanyan illustrates this with a difficult image containing a bird. The model decides that it needs to inspect different portions, increases the contrast, and eventually finds the bird. He connects this behavior to embodied models that understand how to inspect different modalities instead of producing a single fixed box from one pass.

10:45

Embodied reasoning models can orchestrate lower-level tactile policies

For a task such as making coffee, an orchestrator can break the task into subtasks while a tactile control policy handles the lower-level actions. Aghajanyan describes a spectrum between a single end-to-end VLA and a more agentic system with separate reasoning and control components. He also describes using the model for robotic data annotation. The model jumps to different parts of a video, clips the relevant sections, checks whether captions are correct, and self-verifies the annotations. He attributes the practicality of this workflow to the model's speed and lower cost.

08:54

A petabyte-scale multimodal dataset supports the embodied foundation model

Perceptron's released model was trained on a dataset of about one petabyte spanning text, images, videos, and trajectories. The trajectories include desktop use and video games as well as robotics-related data. The training mixture comes from internet crawls, custom training recipes, and synthetic data pipelines. Aghajanyan presents the model as better than Gemini's embodied reasoning model and says it is substantially cheaper, while describing these results as emerging from the combination of perception, reasoning, and control.

12:35

More video pretraining can replace some teleoperation data

Aghajanyan presents scaling behavior for models trained jointly on control, trajectories, perception, and embodied reasoning. Teleoperation data costs about one hundred dollars per hour, so general video pretraining offers a cheaper source of data. Within the amount of compute Perceptron has tested, ten times more video pretraining data can substitute for ten times less teleoperation data. Pure policies still improve with more teleoperation data, but he says the tradeoff is more favorable for unified embodied foundation models.

14:11

One model can read a book title and emit control tokens

A demonstrated policy runs within a single model that performs embodied reasoning and outputs control tokens. In one task, the model reads the title of a book, uses knowledge about what type of book it is, and places it in the appropriate bin. Aghajanyan presents this as a multi-step combination of perception and control that is difficult for a pure end-to-end control model because the model must recognize which perceptive subtasks it needs to complete. He says the policy works zero shot relatively well and that a smaller model was planned for open release in July.

"We want to be able to build physical AI foundations that give us the ability to perceive, understand and interact with the physical world in real time."00:34
Who should watch
  • You are building multimodal or robotic systems and need to decide how perception, reasoning, and control should be combined.
  • Your video or sensor stream is too large for ordinary dense processing, and you want a model to select task-relevant tokens dynamically.
  • You are collecting teleoperation data and want to understand how video pretraining might reduce that cost.