# Vision-Language-Action Models: How a Large Model Gets a Body

How a VLM trained on text and images is wired to motors — action tokens and 256-bin discretization in RT-2, cross-embodiment data in Open X-Embodiment and OpenVLA, and why π0 replaced the token head with a flow-matching action expert.

> vision-language-action models · action chunking · behavior cloning · contact-rich manipulation · About 9 min · Oct 6

## Key points

1. A vision-language-action model is a pretrained vision-language model with an action output interface, a mixed data recipe, and a training schedule — not a new architecture.
2. RT-2's trick was to define the action as eight integers (6-DoF end-effector displacement, gripper, terminate flag), discretize each continuous dimension into 256 bins, and train the VLM to emit them as ordinary tokens, so no new parameters were needed.
3. Co-fine-tuning robot trajectories together with web vision-language data generalizes better than robot-only fine-tuning, because the model keeps its semantic concepts while learning motor commands; models trained from scratch on robot data perform poorly.
4. Web pretraining transfers semantics — object relations, unseen instructions, reasoning about which object to pick — but the model's physical skills stay bounded by the motions present in the robot data.
5. Cross-embodiment training pools many robots into one dataset by coarse alignment: one camera view per dataset, every action converted to a 7-dimensional end-effector vector, normalized before discretization and de-normalized at execution; coordinate frames and absolute/relative/velocity semantics are not aligned, so the same vector means different motions on different robots.
6. Discretizing actions into 256 bins costs resolution (roughly 4 mm per step if an axis spans a meter, 8 bits per dimension) and speed, since a 55B model emits 7–8 tokens per action and is queried only a few times per second.
7. π0 replaces the token head with an action expert: a second set of weights inside the same transformer that predicts a 50-step continuous action chunk via flow matching, starting from noise and refining it toward real actions, which supports control at up to 50 Hz.
8. Even with a learned policy in place, a low-level controller still turns the model's end-effector targets into joint commands and corrects errors, so the closed loop from planning and feedback control has narrowed but not disappeared.
9. Reported generalization gains (3× for RT-2-X, 16.5 absolute points for OpenVLA over a 55B model) come from specific task suites on specific robot platforms and should not be read as open-world competence.

---

The previous article ended with a policy that produces actions in chunks, and with the observation that learned motion and computed motion coexist because contact is unforgiving. This article picks up one specific thread: **what does it actually take to give a model trained on text and images a body?** The word "body" here is not a metaphor. It means a fixed set of numbers leaving a neural network and arriving at motors, and the interesting engineering is in how those numbers are chosen.

## The mismatch: text out, torque in

A vision-language model takes an image and a question and answers in text. It already knows a great deal that a robot needs: that this is a cup and not a bowl, that "the one closest to the sink" picks out one particular object, that a rock can serve as an improvised hammer. What it does not know how to do is emit a joint angle.

So the first question is not "how do we build a brain for robots" — the brain largely exists — but "what is the narrowest possible interface between that brain and an arm?" Everything in this article follows from how the field answered that question.

## Turning actions into words

RT-2, published by Google DeepMind in 2023, gave the simplest answer available: **write the action as text and let the model continue the sentence** [1].

Concretely, the action space is defined as a 6-degree-of-freedom displacement of the end-effector (three position components, three rotation components), plus gripper opening, plus a discrete "task is done" flag. Each continuous dimension is divided into 256 uniform bins, and the bin index is written as an integer. One action therefore becomes eight numbers, and a training target looks like this:

```
1 128 91 241 5 101 127
```

The model's existing tokenizer is then told that 256 of its tokens stand for those bins, and training proceeds with the ordinary next-token objective. At inference on a robot task, decoding is restricted so the model can only emit valid action tokens. No new parameters, no new architecture — a pre-existing VLM, fine-tuned to output text-encoded actions. That is the whole recipe, and the authors named the category **vision-language-action (VLA) models**.

There is a detail in that recipe worth pausing on, because it is the part that makes the trick work rather than merely function. You can fine-tune a VLM on robot data alone, and you get a policy that is a bit better than nothing. RT-2 was instead **co-fine-tuned**: robot trajectories are mixed into the training batches together with the original web-scale vision-language data (roughly one-to-one, with the robot data upweighted). The authors report that co-fine-tuning generalizes better than plain robot-only fine-tuning, and plausibly why: the model keeps seeing abstract visual concepts while it learns motor commands, so it does not overwrite its semantic knowledge to fit a small robot dataset. Two control experiments support this reading. Training a comparable model from scratch — no VLM weights at all — performs badly even at 5B parameters. And within the co-fine-tuning setup, the 55B version generalizes better than the 5B version.

The payoff is specific, and the boundary is just as specific. The model can place an object on a number or an icon it never saw in robot demonstrations, pick the smallest object, or, with a chain-of-thought prompt, decide that a rock is what you would use as an improvised hammer. But the paper is explicit that this is semantic transfer, not motor transfer: the physical skills stay inside the distribution of motions present in the robot data. The internet teaches the model what to do; the robot data alone decides what the arm can physically do.

![OpenVLA architecture: a vision encoder and projector feed a language model backbone that outputs tokenized actions](https://arxiv.org/html/2406.09246v3/openvla_model.png)

## One body is not enough

Robot data is the scarce resource. Hours of demonstrations, collected by teleoperation, is the order of magnitude; internet text is not. The obvious move is to pool data across labs and robots, and that ran into a formatting problem. Different robots have different numbers of joints, different camera placements, different control interfaces.

Open X-Embodiment attacked this by **coarse alignment** [2]. Its dataset pools 60 existing datasets from 34 labs into 1M+ trajectories covering 22 robot embodiments — single arms, bimanual robots, quadrupeds — and 527 skills. To make the data usable in one model, each dataset contributes one canonical camera view, resized to a common resolution, and every action is converted into the same 7-dimensional end-effector vector: $x, y, z$, roll, pitch, yaw, gripper. Each dataset's actions are normalized before discretization and de-normalized when executed, so the same model output can mean different physical displacements on different robots.

Note what is *not* aligned: coordinate frames, and whether a number is an absolute position, a relative displacement, or a velocity. Those come from whatever control scheme each robot originally used. The consequence is blunt — the same action vector can induce very different motions on different robots, and camera poses vary too. The model is not being handed a clean universal motor language; it has to read the image and the context to interpret its own output. Despite that, pooling helped: RT-1-X trained on the pooled mixture outperformed the original methods of the contributing labs by a reported 50%, and the VLM-based RT-2-X improved generalization roughly threefold over a model trained only on the evaluation robot's data.

Data diversity also turned out to substitute for scale. OpenVLA, a 7B-parameter open-source VLA trained on 970k trajectories from that same collection, reported beating the 55B RT-2-X by 16.5 percentage points of absolute success rate across 29 tasks on two robot platforms, with the fused DINOv2 + SigLIP vision encoder feeding a Llama 2 backbone and actions discretized into 256 bins per dimension [3]. One implementation difference is worth knowing because it recurs: instead of dividing the min-to-max range, OpenVLA sets bin widths from the 1st to 99th quantile of the training actions, so a few outlier actions cannot stretch the range and coarsen every bin.

## The price of writing actions as text

Discretization buys compatibility with a language model, and it costs two things: rate and resolution.

Resolution first, with a rough estimate: if one axis of a task spans about a meter and you give it 256 uniform bins, one bin step is roughly 4 mm. Each dimension also only carries $\log_2 256 = 8$ bits, so a 7-dimensional action carries 56 bits. That is the entire motor bandwidth of a 55-billion-parameter model. For pushing blocks around, 4 mm is usually fine. For inserting a peg or folding cloth, it is not.

Rate is the sharper constraint. Pooled datasets like Open X-Embodiment assume a policy queried a few times per second — the reference inference setup runs at about 3 Hz, meaning one query every 333 milliseconds. A 55B model generates its action tokens one at a time, autoregressively. Seven to eight tokens per action, plus the latency of a large forward pass, gives you a handful of control updates per second. Queries can be batched and served on accelerators to keep up, but the ceiling is structural: the more tokens you need, the slower the loop, and dexterous manipulation wants corrections far faster than a few hertz.

## A different answer: continuous chunks from a continuous head

The alternative is to stop making the action head speak in tokens at all. π0 (pi-zero) keeps the VLM and replaces the output mechanism [4].

Its backbone is PaliGemma, a 3B vision-language model. Images and language go through the backbone's weights as usual. Actions and the robot's own joint state go through a second, smaller set of weights — around 300M parameters — that the authors call the **action expert**, with everything attending over one shared sequence inside a single transformer. Functionally it is a mixture of experts with two members: one specialized for seeing and reading, one specialized for moving.

That head does not emit one action at a time. It emits a whole **chunk** of $H = 50$ future actions, and it emits them as continuous values using **flow matching**. The training idea is straightforward once you picture it. Take a real chunk of 50 actions from the data. Mix it with pure random noise using an interpolation parameter $\tau$, so at $\tau = 1$ you have the real actions and at $\tau = 0$ you have noise. Ask the network to predict, at whatever $\tau$ it was given, the direction that points from the noise toward the real actions. At inference you reverse the process: start from noise and integrate that predicted direction over a few steps, conditioned on the current images, instruction and joint state, and you get a continuous action chunk.

Why this matters for the two problems above: a chunk of 50 actions arrives in one generation, not 350 tokens of autoregressive decoding, and the values are not snapped to 256 levels. π0 is trained on roughly 10,000 hours of demonstrations across single-arm, dual-arm and mobile manipulators plus the open cross-embodiment data, with a pretraining stage on the broad mixture followed by a post-training stage on narrower, higher-quality data to induce dexterity and speed. It reports controlling robots at up to 50 Hz on tasks like folding laundry and assembling boxes.

## What a VLA is, and what it is not

Pulling the thread together, a VLA in this sense is not a new kind of brain. It is a pretrained vision-language model plus an **action interface** — tokenized bins or a continuous head — plus a **data mixture** — robot trajectories, cross-embodiment data, and leftover web data — plus a training schedule. "Getting a body" means adding an output head, a loss function and a data pipeline, not simulating a nervous system.

Three boundaries are worth keeping attached to the claim:

- **The loop from the previous article still exists.** A VLA outputs an end-effector displacement or a chunk of them; a low-level controller still converts that into joint commands at kilohertz rates and still corrects for the fact that the model's prediction is slightly wrong. Measured at 50 Hz, the learned policy has moved down a level, but it has not replaced control.
- **Semantics transfer; skill does not.** Every paper here says some version of the same thing: web pretraining transfers nouns, relations and reasoning, and the physical repertoire stays bounded by the robot data. If folding cloth is absent from the demonstrations, no amount of internet text produces the motion.
- **The generalization numbers are benchmark numbers.** "3× better" or "16.5 points better" refer to specific task suites on specific robots. They are real measurements, and they are not claims about open-world competence.

The strong version of the popular story — that robots are simply next in line for the recipe that produced large language models — is half right. The scaling advice carries over: bigger backbones generalize better, models trained from scratch on robot data fail, and data diversity can beat parameter count. What does not carry over is the interface. Language models emit sentences that a human can read and check; a VLA emits 56 bits of displacement into a physical world with friction, contact, and no undo. The work of the last few years has been mostly about that interface, and it is still where the difficulty lives.

<details>
<summary>Check yourself</summary>

**Why can a model fine-tuned only on robot trajectories perform worse than one co-trained on web data?**
Because robot-only fine-tuning tends to overwrite the semantic knowledge learned from web-scale data, while mixing web data into the batches keeps those concepts alive during motor learning. RT-2's ablations show co-fine-tuning generalizing better than robot-only fine-tuning, and a from-scratch model performing badly.

**If binning an action into 256 levels is lossy, why did RT-2 do it?**
Because bins let actions become ordinary tokens, which lets an existing VLM's next-token machinery produce actions without new parameters or a new architecture — trading resolution and inference speed for reuse of pretrained weights.

**π0 also uses a VLM. What exactly changed?**
The output head and its training objective. Instead of predicting one discrete action token at a time, a smaller action expert predicts a 50-step continuous action chunk via flow matching, starting from noise and refining it toward a real chunk — which raises both resolution and control frequency.
</details>

## Sources

1. [RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control — action-as-text-token recipe, 256-bin discretization, co-fine-tuning ablations, emergent semantic capabilities](https://arxiv.org/abs/2307.15818)
2. [Open X-Embodiment: Robotic Learning Datasets and RT-X Models — 1M+ trajectories across 22 embodiments, coarse action alignment and normalization, RT-1-X and RT-2-X results](https://arxiv.org/abs/2310.08864)
3. [OpenVLA: An Open-Source Vision-Language-Action Model — 7B model trained on 970k trajectories, quantile-based action binning, comparison against 55B RT-2-X](https://arxiv.org/abs/2406.09246)
4. [π0: A Vision-Language-Action Flow Model for General Robot Control — PaliGemma backbone, action expert, flow matching over 50-step action chunks, up to 50 Hz control](https://arxiv.org/abs/2410.24164)

---

Original article: https://eulore.ai/articles/vision-language-action-models-action-interface-caf07f6d

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
