A robot has no eyes in the human sense. When a camera points at a kitchen counter, the robot receives a grid of numbers — brightness in red, green and blue for each of, say, 1920×1080 positions. Nothing in that array says counter, knife, or two meters away. The distance information simply isn't there. Yet a few hundred milliseconds later, some piece of software asserts that there is an obstacle 60 cm ahead and slightly left, and the arm moves accordingly.

The gap between those two moments is what this article is about. It is one of the central problems of embodied AI: turning raw, noisy sensor streams into an internal picture of the world that a robot can act on. In the last article we looked at why physical skills resist automation; here we look at the perception side — what a robot must infer before it can move at all.

What a sensor actually hands you

Different sensors fail in different ways, and the pattern of failure shapes everything downstream.

A camera measures light. It gives dense, rich appearance information — color, texture, edges — but no distance. Recovering depth requires doing extra work: two cameras and triangulation, a projector that paints a pattern onto the scene, or a learned model that guesses depth from a single image. All of these produce depth estimates that are noisy, and worse, noisy in ways that depend on the scene (reflective floors, blank walls, direct sunlight).

A lidar measures distance directly: it fires a laser, times the return, and reports a range plus the bearing it was pointed at. So it gives you a point in 3D space immediately — but only for the thin slice of the world the beam hit. It says nothing about what the surface is. A black plastic bag and a wall look identical in a point cloud.

Proprioceptive sensors tell the robot about itself, not the world: joint encoders measure the angle of each motor, an inertial measurement unit (IMU) measures acceleration and rotation rate. And force/torque sensors measure contact, which is what a robot needs when it can no longer see the thing it's holding.

Two properties are shared by all of them. First, every reading is expressed relative to the sensor, not the world: "an object 1.2 m away at 30° to my left." Second, every reading is partial — a camera sees what it faces and nothing behind the counter; lidar sees one plane per sweep. Sensing gives you a limited, sensor-centric, noisy slice of reality.

Why the same number means different things in different places

Here is a step that looks like bookkeeping and is actually the crux.

To place a lidar return into a map — "there is a wall here, in the kitchen, next to the fridge" — you must convert from the sensor's frame to the robot's frame, and then from the robot's frame to the world frame. The first conversion is fixed geometry: you know where the lidar is bolted onto the chassis. The second is not fixed at all. It is where the robot is, and the robot has to figure that out.

This is why the chain world ← robot ← sensor is the hinge of the whole problem. A range reading is only interpretable as evidence about the world if you already have a hypothesis about where you are standing. Get the pose wrong by 30 cm and you will carve your hallway 30 cm off from the real one, then carve it further off on the next sweep.

A robot can estimate how far it has moved by counting wheel rotations or by integrating the IMU. This is odometry, and it drifts. Each tiny error in wheel slip or gyro bias is added to the last one, so the error grows without bound: over a corridor, centimeters; over a few minutes, meters. Odometry alone can never anchor you to the world, because nothing in it ever refers to anything outside the robot.

Perception is inference, not reading

The standard way to formalize all this is to admit that the robot never observes the state of the world. It observes evidence about it, and maintains a belief: a probability distribution over the things it cares about — its own pose, the location of obstacles, the position of objects.

The mechanism is a two-step recursion repeated at every time step. In words: predict, then correct.

bel(xt)=η p(zt∣xt)∫p(xt∣xt−1,ut) bel(xt−1) dxt−1⏟prediction\text{bel}(x_t) = \eta \, p(z_t \mid x_t) \underbrace{\int p(x_t \mid x_{t-1}, u_t)\,\text{bel}(x_{t-1})\,dx_{t-1}}_{\text{prediction}}

The right-hand integral is the prediction step. It uses a motion model — a description of how the state tends to change given the command the robot just issued (utu_t). Drive forward one meter and the belief shifts forward, and also spreads out, because the motion is not exact.

The left-hand factor is the correction step. It uses a sensor model, p(zt∣xt)p(z_t \mid x_t): how likely is the reading I actually got, if the state were this particular hypothesis? A hypothesis that predicts the reading well gets boosted; one that predicts it poorly gets suppressed. The constant η\eta just rescales everything back into a probability distribution.

This is Bayes' rule with a habit. What makes it a filter rather than a one-shot calculation is that the belief from the previous step becomes the input to the next prediction, so evidence accumulates over time.

The arithmetic of a map, concretely

The occupancy grid, developed in the late 1980s by Alberto Elfes and Hans Moravec, is this idea applied to space itself: divide the floor plan into cells, and let each cell carry a probability that it is occupied rather than empty 12.

Cells update by odds, where odds=p/(1−p)\text{odds} = p/(1-p), because the update rule is just multiplication:

p(occupied∣all readings)p(empty∣all readings)=prior odds×∏ip(zi∣occupied)p(zi∣empty)\frac{p(\text{occupied} \mid \text{all readings})}{p(\text{empty} \mid \text{all readings})} = \text{prior odds} \times \prod_i \frac{p(z_i \mid \text{occupied})}{p(z_i \mid \text{empty})}

Suppose the sensor model says a "hit" return is 5 times more likely if the cell is occupied than if it is empty, so each hit contributes a factor of 5. Start with the honest prior that you know nothing: p=0.5p = 0.5, odds =1= 1.

  • One hit: odds =1×5=5= 1 \times 5 = 5, so p=5/6≈0.83p = 5/6 \approx 0.83.
  • A second hit from a different vantage point: odds =25= 25, so p=25/26≈0.96p = 25/26 \approx 0.96.
  • Then a reading that passes straight through that cell, contributing a factor of 1/51/5: odds =5= 5, back to p≈0.83p \approx 0.83.

Two things are worth noticing. First, if you work in log-odds instead, the multiplication becomes addition: log⁡5≈1.61\log 5 \approx 1.61 gets added per hit and subtracted per miss. This is why this update is cheap enough to run over a million cells in real time. Second, no single reading decides anything. The map is a running account of accumulated weak evidence, and it is never more certain than the arithmetic says it should be.

That last point is where the caveats live. The formula above treats readings as independent, but they usually aren't: one reflective window can produce false returns in many lidar beams at once, and the grid will grow confidently wrong. It also assumes cells are independent, which is false whenever real objects cover several of them; Elfes noted this explicitly and treated it as a simplifying assumption 1. And real sensors are messy in sensor-specific ways — Moravec's early work used a ring of 24 Polaroid sonar transducers, and sonar has large lateral uncertainty (it's vague about the angle) but small range uncertainty, exactly the inverse of a stereo camera's error profile 2. The sensor model has to encode that.

The loop with no starting point

Now the two halves collide. To write a reading into the map, you need your pose; to know your pose, you need the map. Each depends on the other, and neither is given.

This is SLAM — simultaneous localization and mapping — and it is worth being precise about why the obvious fix fails. You might try to solve the two problems separately: build a map assuming your poses are right, then localize in the map you built. Durrant-Whyte and Bailey's tutorial points out that the joint probability over poses and landmarks cannot be factored into the product of a pose term and a map term; assuming that it can produces estimates that are inconsistent — that is, the filter's stated uncertainty is smaller than its real error 3.

The reason is that the errors are not independent. Almost all of the error in your estimate of any given wall comes from one shared source: not knowing exactly where the robot was when it saw the wall. So the walls are misplaced together, which means the geometry between them can be very accurate even when their absolute positions have drifted — the tutorial notes that the relative position of two landmarks may be known precisely while the absolute position of each is quite uncertain 3. That single observation is what makes SLAM solvable at all, and it's why loop closure matters: when the robot returns to a place it recognizes, it can suddenly pin down everything it saw in between.

Practically, SLAM is implemented with filters — the extended Kalman filter, which tracks everything as Gaussians, or Rao-Blackwellized particle filters, which sample many candidate trajectories (FastSLAM) 3. Durrant-Whyte and Bailey judged the problem solved at a theoretical level by 2006, while noting that substantial gaps remained in making the maps perceptually rich 3. That gap is real: a grid tells you a cell is occupied, not that it is a chair, and not that chairs are for sitting on.

Perception is something the robot does

There is a second reason the picture above is incomplete: it treats the robot as a passive observer accumulating evidence. But the robot has a body, and it can move the sensor.

Ruzena Bajcsy made this argument in 1988 under the name active perception: "we do not just see, we look" 4. Her point was that perception is a control problem — you choose where to point the camera, when to move closer, what to look for in the data — precisely because the data you'd get is a function of what you do. A later retrospective sharpens the test: an agent is an active perceiver if it knows why it wants to sense, and then chooses what, how, when and where to perceive 5.

This closes the loop in both directions. Actions change the world, which changes what the sensors will report next; and uncertainty is itself a reason to act. A robot that is unsure whether the space ahead is clear can resolve that uncertainty by moving and looking, and this is a legitimate — sometimes optimal — use of motion.

It also means the map is not a fixed product. An occupancy grid is one choice among several: a 3D voxel grid, a sparse set of landmarks, a point cloud, or a layered arrangement with a geometric layer for collision and a semantic layer for meaning. Elfes framed the occupancy grid as a reaction against what he called the geometric paradigm, which forced early, brittle decisions about what the sensor data meant; keeping probability in every cell defers those decisions until more evidence arrives 1.

What the large models changed — and didn't

Vision-language-action (VLA) models mostly sidestep the explicit map. The pattern, set by RT-2, extends a pretrained vision-language model with robot data and casts the robot's actions as text tokens, so the same training corpus covers "describe this image" and "move the gripper here" 6.

That buys semantics for free. Web-scale pretraining supplies object knowledge and language grounding that a hand-built occupancy grid never had, which is why such policies generalize to instructions and objects they were never trained on. What it does not obviously buy is metric precision. Survey work in this area notes that representations learned for image-text tasks are not automatically suited to control, and that auxiliary objectives — segmenting a target object, predicting keypoints or contact points, estimating depth — generally help 7. Several systems reinsert explicit 3D structure for exactly this reason: PerAct operates on voxel maps built from RGB-D, Act3D uses a continuous 3D feature field, and RVT re-renders views to get 3D-like inputs back without voxelizing 8.

The pattern is not that explicit maps were wrong and learned features replaced them. It's that the two do different jobs. Grids and filters are strong at accumulating geometry and tracking uncertainty over time, and weak at knowing what anything is. Pretrained multimodal models are the reverse. Where the boundary settles is still open.

The short version

A robot's "understanding" of a scene is not a picture of the world but a maintained belief about it. Sensor readings are noisy, sensor-centric, partial evidence. Turning them into something actionable requires two models — how the world changes when you act, and how the world would look if a given hypothesis were true — combined by repeated prediction and correction. And because placing any reading requires knowing where you were when you took it, localization and mapping have to be solved together, not one after the other. Perception is not something that happens to a robot before it acts; it is part of the acting.