The last article left a robot with a maintained belief about the world — a probability distribution over where things are, updated every time a new sensor reading arrives. That belief has not done anything yet. A robot is not a camera with opinions; it is a machine that changes the physical world, and the interesting part is that changing the world changes what it will perceive next. Action closes a loop through physics, and physics does not forgive: time runs one direction, an arm that pushes a glass off a table cannot un-push it.
So the question this article is organized around is narrow and concrete: given some belief about the world, where do the numbers that go to the motors come from? There are two broad families of answers, and the shift between them is the biggest story in robotics over the last decade. One family computes the motion at runtime. The other learns it in advance. They are not enemies, and understanding why they coexist is most of what "how robots act" means today.
What a motion actually is
Before the two families, one piece of vocabulary. A robot arm with 7 joints plus a gripper needs 8 numbers to describe its pose completely. That collection of numbers is its configuration, and a motion is a path through that 8-dimensional space. Actuators can be commanded at different levels of that space: directly in torque (most raw, fastest loop required), in joint angles (a low-level controller holds the angle for you), or in the end-effector frame ("move the hand 2 cm forward and rotate it 15°"). Learned policies almost always command the third kind, and we will come back to why that matters.
Concretely: RT-2, a widely cited vision-language-action model, outputs 6 numbers for the hand's positional and rotational displacement, plus one for gripper opening, plus a flag meaning "the task is done" 1. That is the entire interface between a 55-billion-parameter neural network and a physical arm — 8 numbers, each quantized into 256 bins and written out as a string of integers.
Family one: compute the motion
The classical approach decomposes acting into two jobs. The first is planning: with a model of how the robot moves, search for a sequence of commands that reaches the goal. The model is usually written as
where is the state, is the command, and encodes the physics — how joint torques turn into accelerations, how wheels roll, how objects move when pushed. Planning means choosing a sequence that minimizes a cost such as "distance to goal, plus energy, plus penalty for hitting things." The output is a trajectory, and it is computed, by a numerical optimizer, on the robot, at runtime.
The second job is control, and it exists because the executed motion is always slightly wrong. The real robot's mass is not exactly the model's mass, the floor tilts, the payload is heavier than assumed. So the controller does not just replay the plan; it measures the difference between where the robot is and where the plan says it should be, and pushes back harder when the error is bigger. That is the core of feedback control: a response proportional to the error, usually plus a term that reacts to how fast the error is changing (to avoid overshooting) and one that slowly accumulates (to erase a persistent offset). Feedback is the reason a robot can be ignorant about many details of its own physics — the loop absorbs them.
The modern version of "compute at runtime" merges both jobs and is called model predictive control (MPC): solve the optimization over a short horizon, execute only the first command, then throw the plan away and re-solve with fresh measurements. It runs tens of times per second. If that sounds like replanning every step, it is — and remember it, because the learned approach will reinvent it.
Where computed motion breaks down
For a large class of problems, this works spectacularly well. Welding robots, pick-and-place machines, and quadrupeds doing dynamic gaits have all been driven by MPC that reasons explicitly about contact forces and footholds 7.
The problem is what happens when the robot must decide when and whether to touch things. Free-space motion is a geometry problem; contact is a logic problem glued to a physics problem. If the foot is on the ground, the equations of motion take one form; if it is in the air, they take another; if it is sliding, a third. Planning must therefore choose both a continuous trajectory and a discrete sequence of contact modes, and the number of candidate sequences grows combinatorially. Standard formulations turn this into a mixed-integer optimization whose runtime balloons with the number of possible contacts, which is why multi-contact MPC has historically been too slow for real-time use 6. And even when you solve it, the answer depends on the friction coefficient between two surfaces you measured once, badly.
So the classical stack is powerful exactly where a decent model can be written down, and strained exactly where the interesting physical interactions live: grasping, slipping, deformable objects, tools, cluttered shelves.
Family two: learn the motion
The learned approach replaces the entire runtime optimization with a single function:
a policy that maps an observation (camera images, joint angles, a language instruction) to a distribution over actions . The right way to read this is as a shift in where the optimization happens. In the classical stack, a cost function and a physics model are optimized on the robot, every 50 milliseconds. In a policy, the result of that optimization has already been baked into the network's weights during training. There is no model and no cost function present at runtime at all — only a very large lookup function that was shaped by one.
There are two ways to shape it. Imitation: collect demonstrations of a human or a scripted expert doing the task, and train the network to reproduce the demonstrated action for the observed situation. Reinforcement learning: let the robot try, score the outcomes, and make the actions that led to good outcomes more likely. Imitation is far more common in current manipulation systems, for the boring reason that it needs a person with a joystick rather than millions of real-world trials.
Why the actions come in chunks
Here a purely practical constraint turns into a design principle. A policy has to run in the control loop, and large models are slow. RT-2's largest version — 55B parameters — could only be queried at 1–3 Hz, and even its 5B version at around 5 Hz 1. But smooth manipulation wants commands at tens of hertz. A robot that recomputes a single action every 300 ms will twitch.
The fix is action chunking: predict a whole sequence of future actions in one forward pass, then execute them without re-querying the model. A 50-action chunk at 50 Hz buys a full second of motion per forward pass. π0, a vision-language-action model from Physical Intelligence, uses exactly this — a chunk of future actions generated by a flow-matching head, driving robots at up to 50 Hz on tasks like folding laundry 3. Diffusion Policy, an influential earlier method, does the same thing with a denoising diffusion model and explicitly borrows the classical habit of re-planning before the chunk runs out 4.
Chunking also fixes a deeper problem that has haunted imitation learning from the beginning: compounding error. A policy is trained on states that the expert visited. At deployment it makes a small mistake, and that mistake carries it into a state the expert never visited, where the policy is even less accurate, which produces a bigger mistake, and so on. The classical analysis bounds this: if the policy errs with probability per step, behavior cloning's worst-case regret grows as — quadratically in the length of the task — while methods that let the learner practice with an expert correcting it achieve the optimal 8. That gap is why imitation learning often looks excellent in held-out-action accuracy and then fails on the robot.
Chunking attacks this by making the loop more forgiving: for several steps, the robot acts on the same observation, so an error does not immediately feed back into a fresh, wrong prediction. A control-theoretic analysis shows that the required chunk length grows only logarithmically in the system's stability parameters, and — importantly — that this benefit depends on the underlying system being stable in open loop, which in manipulation is exactly why these policies command the end-effector and let a classical low-level controller handle the servo loop 5. A recent study argues the mechanism has more parts than that, crediting both the reduction in compounding error and an ensemble-like effect from learning many temporal relationships at once 9. Treat the precise credit split as unresolved; treat the empirical benefit as solid.
So the learned policy is not floating free of control theory. It sits on top of a classical feedback controller, and it behaves like a receding-horizon controller whose model has been absorbed into its weights.
What changed, and what didn't
What genuinely changed is generalization. Because these models are built on vision-language backbones pretrained on internet data, they inherit semantic knowledge that no hand-written planner has. RT-2 could follow instructions never seen in robot data — placing an object on a specific number, or picking the object best used as an improvised hammer — and transferred web-scale concepts to motor behavior 1. OpenVLA, a 7B open-source model fine-tuned on 970k real demonstrations, outperformed a 55B predecessor on generalist manipulation with seven times fewer parameters, largely by collecting more diverse data 2. π0 was trained on roughly 10,000 hours of demonstrations across several robot bodies 3. The recipe is now recognizably the language-model recipe: large diverse pretraining, then careful post-training. π0's authors note a detail worth remembering: training only on high-quality demonstrations teaches a policy to perform tasks well but never to recover from mistakes, because mistakes almost never appear in such data.
What did not change is that these policies have no guarantees. They are functions that interpolate over the distribution of what they were shown; outside it, they degrade, silently and confidently. They still need a low-level controller to hold the arm steady between chunks. And they still cannot, in any meaningful sense, tell you why they moved — the physics that the classical planner wrote in the open is now distributed across billions of weights.
That is the honest state of the field: a runtime optimizer that knows its model is wrong, and a learned function that never had one. The current best practice is usually to put them in the same stack, with the learned part handling the contact-rich mess and the computed part handling everything that still has a clean answer.
Check yourself
Why does action chunking help a slow policy? Because one forward pass produces actions instead of one, so the policy can be queried rarely while the robot still receives commands at high frequency. A 50-action chunk at 50 Hz covers a second of motion per query.
Why does chunking reduce compounding error, intuitively? For several steps the robot keeps acting on the same observation, so a small deviation cannot immediately feed back into a new, differently wrong prediction. The error no longer amplifies step by step.
What did MPC and chunked policies end up sharing? Both plan over a short horizon, execute part of it, and re-plan — MPC using a physics model solved on the fly, chunked policies using a learned function whose training already absorbed that work.