# Moravec's Paradox: Why the Easy Things Are Hardest for Robots

Why chess fell to machines before walking did — the evolutionary story, the engineering account behind it, and what vision-language-action models have and haven't changed.

> Moravec's paradox · contact-rich manipulation · sim-to-real gap · vision-language-action models · About 10 min · Oct 6

## Key points

1. Moravec's paradox is the observation, stated in Hans Moravec's 1988 book Mind Children, that computers achieved adult-level performance on intelligence tests and checkers while failing to match a one-year-old's perception and mobility.
2. It reads as a paradox only because we judge difficulty with two unreliable rulers: how much better the best human is than the median human (variance), and how much conscious effort a task takes us.
3. Effort is a poor difficulty signal because fluent, well-optimized mental processes leave no trace in awareness — as Minsky put it, we are least aware of what our minds do best.
4. Moravec's own explanation is evolutionary: skills that have been shaped by natural selection for longer should be harder to reverse-engineer, and the oldest skills are largely unconscious. This story is intuitive but contested and is not empirically settled.
5. A separate engineering account explains the same reversal without claims about brains: task difficulty for a learning system scales with the size of the search space and with how sparse and how expensive the feedback is.
6. Chess has a branching factor of roughly 35, exact rules, a perfect simulator and unlimited free self-play; manipulation has continuous high-dimensional actions, discontinuous contact dynamics, sparse rewards and a practice budget of a few attempts per minute.
7. Contact is the crux: s rigid-body physics simulates acceptably, but friction, cloth, cables and deformation are where simulation and reality diverge, so the sim-to-real gap grows exactly where Moravec's paradox points.
8. On tasks like peg-in-hole and USB insertion, sparse-reward reinforcement learning can fail to achieve any success within the training budget, while humans treat the same task as trivial.
9. Vision-language-action models such as RT-2 transfer semantic generalization from web-scale data to robot control, but the RT-2 authors state the robot gains no new motions — physical skills stay limited to the distribution seen in robot data.
10. Evaluation numbers fit the pattern: a VLA comparison found about 72% in-distribution success for π0, with failures dominated by grasping and releasing rather than command understanding, and first-step pick-up success dropping from roughly 75% to 35% when a two-step instruction was given.
11. Control frequency is a second wall: contact-rich tasks want tens of hertz, while the largest VLAs run in the single-digit hertz range when the language model is inside the control loop.
12. When reading a robot demo, ask whether it shows semantic generalization (the foundation-model side) or the last centimeter of contact and sensing (the part Moravec flagged), and check success rates across new positions rather than a single successful take.

---

In 1988 the roboticist Hans Moravec wrote a sentence that has aged uncomfortably well:

> It is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility. [1]

This is Moravec's paradox. It is not really a paradox — it is a reversal. We expect difficulty to line up between humans and machines, and it doesn't. The things that took you years of schooling are things machines learned first; the things you do before breakfast are the ones still unsolved.

For anyone following robotics news, this reframing is worth understanding, because it tells you where to look when a demo looks impressive and where the work actually is.

## What the claim actually says

Three clarifications, because the slogan gets flattened into "robots are bad at easy stuff."

First, it is a claim about ranking, not absolute difficulty. Moravec isn't saying walking is harder than chess in some cosmic sense. He's saying the *order* in which these problems yield to engineering is roughly the reverse of the order in which they feel hard to us.

Second, it was not one person's idea alone. Moravec, Rodney Brooks and Marvin Minsky converged on it in the 1980s. Brooks pointed out that early AI had quietly defined intelligence as "the things that highly educated male scientists found challenging" — chess, symbolic integration, theorem proving — while ignoring "the things that children of four or five years could do effortlessly," like telling a coffee cup from a chair or finding the way from the bedroom to the living room. Minsky added the diagnostic clue: "In general, we're least aware of what our minds do best." Steven Pinker compressed it in 1994: "the main lesson of thirty-five years of AI research is that the hard problems are easy and the easy problems are hard." [1]

Third, it is a tendency with exceptions, not a law. One of its original flagship examples — computer vision — largely fell to deep learning around 2012 once GPUs made large convolutional networks trainable. So the paradox is best used as a heuristic about where difficulty hides, not as a prediction that certain problems are impossible. [2]

## Why our sense of "hard" is misleading

The reversal only looks strange because we use ourselves as the measuring stick, and we use two bad rulers.

The first is *variance*. Ask how much better the best human is than a median human. In mathematics or chess, the gap is enormous — a grandmaster beats a casual player every time and there is a clear ladder of skill. In walking or catching a ball, the gap is tiny; nearly every adult does it about equally well. We read low variance as "easy." But low variance across a whole species is often a sign of something that has been optimized very hard and very long, until essentially everybody can do it.

The second bad ruler is *conscious effort*. Introspection only shows us the parts of our thinking that run badly. Fluent, well-optimized processes generate no felt effort and leave no trace in awareness — which is exactly why they seem trivial. Effort is a signal of poor optimization, not of difficulty.

## Moravec's own explanation, and its limits

Moravec offered an evolutionary answer, which reduces to three steps: the difficulty of reverse-engineering a skill should be roughly proportional to how long that skill has been evolving in animals; the oldest skills are largely unconscious and therefore feel effortless; therefore the skills that feel effortless should be the hardest to rebuild.

His phrasing: encoded in our "large, highly evolved sensory and motor portions" is "a billion years of experience about the nature of the world and how to survive in it," while "the deliberate process we call reasoning" is "the thinnest veneer of human thought." Abstract thought is "a new trick, perhaps less than 100 thousand years old." [1]

That story is intuitive, and it is also contested. A careful fact-check argues it is not empirically backed and that the evolutionary part is speculative — AI researchers have a habit of making up stories about brains. Two of its predictions have also wobbled. Reasoning in open-ended settings turned out to be genuinely hard for AI: symbolic planners that worked in chess were brittle in the messy real world, and common sense — supposedly an "easy for humans" skill — remains a bottleneck. And vision, once the headline example of an easy-for-humans problem, stopped being one. [2]

So Moravec's evolutionary argument is better treated as a hypothesis about *why* the reversal exists than as a proven mechanism. Fortunately, there is a second explanation that doesn't depend on claims about brains at all, and it is the one that actually explains what roboticists struggle with.

## The engineering account: search space, feedback, and cheap practice

Strip away humans entirely and ask what makes a problem hard for *any* learning system. Two properties do most of the work.

**How big is the search space?** At each step, how many actions are available, over what range, and how many steps before the task is done?

Chess is modest here: roughly 35 legal moves in a typical position, and a game lasting on the order of 40 moves. The state space is astronomically large in the abstract, but the branching is small and the rules are exact.

A robot arm is the opposite. Seven joints, each commanded tens to hundreds of times per second, with continuous values rather than a discrete menu; a multi-fingered hand pushes that to twenty-plus degrees of freedom; and the robot's own body, the object, and the surface all interact. Contact is the killer: it makes the dynamics discontinuous. A millimeter of difference in approach angle can flip the outcome from sliding to sticking to jamming. There is no exact rulebook for that; the "rules" are whatever the physics does.

**How dense and how cheap is the feedback?** Chess engines get a reward at the end of every game and can play millions of games per second against themselves. A robot attempting to fold a shirt gets success or failure after hundreds of steps, spends seconds per attempt, and each attempt risks damaging hardware or surroundings. Real-world reinforcement learning is further constrained because exploration by definition tries things that might be unsafe. [6]

Now put the two together. It isn't that chess is "more intelligent" than walking. It's that chess is a small branching search problem with a perfect simulator and unlimited free practice, while manipulation is a high-dimensional search problem with discontinuous dynamics, sparse rewards, a simulator that doesn't match reality, and a practice budget measured in attempts per minute. [3]

The simulator gap deserves its own sentence, because it is where the paradox bites hardest. Rigid-body physics simulates reasonably well; contact, friction, cloth, cables and deformation are exactly where simulation diverges from reality. So the sim-to-real gap *grows* with contact complexity — precisely the regime Moravec's paradox points at. That is why real-to-sim tuning, domain randomization, and online adaptation exist: they are workarounds for a physics model that is wrong in the places that matter. [6]

## The peg and the hole

Take the most mundane task imaginable: inserting a USB plug. A human does it in about two seconds, with a wiggle, using force feedback they never consciously process.

For a robot this decomposes into: estimate the socket pose to within roughly a millimeter; maintain a compliant contact rather than a rigid one, or the parts jam and the wrist takes the load; and detect the contact state — sliding, wedged, seated — which is visible in force and torque signals but not in any camera image. Classical control works here, but only if you can write down a contact model that captures the nonlinear, jumpy behavior of two surfaces meeting, and that model has to be re-derived for every clearance and material.

Learning-based approaches sidestep the model, and then hit the data problem. In published comparisons on peg-in-hole and USB insertion, a policy trained from scratch with a sparse success/failure reward "does not achieve any success within the training budget" — thousands of training steps with no signal to climb. [7] Meanwhile the human who has spent a lifetime handling objects makes it look like nothing.

This is what Moravec's paradox feels like in a lab: not a philosophical curiosity but an engineering wall made of contact, tolerances, and the cost of real-world practice.

## Does the large-model era dissolve the paradox?

Partly, and the way it has played out is instructive.

Language turned out to be the opposite of a Moravec-hard problem. Text comes in enormous quantities, every next token provides a training signal, and the meaningful per-step search space is small once you have a reasonable model. That is a favourable regime on both axes — dense feedback, cheap data — and large language models exploited it.

Vision-language-action (VLA) models reuse that machinery for robots. RT-2, for example, encodes robot actions as text tokens and co-fine-tunes a vision-language model on both web data and robot trajectories, which produces real generalization: it follows instructions about objects and concepts that never appeared in robot training data, and can do a little reasoning about which object to pick up. [4]

But the RT-2 paper is explicit about what does *not* transfer: the robot "does not acquire any ability to perform new motions by virtue of including this additional experience." Its physical skills stay limited to the distribution of motions in the robot demonstration data; the model learns to deploy existing skills in new ways. Semantic generalization improved; the motor repertoire did not. [4]

Benchmark numbers show the same shape. A study evaluating several VLA models on manipulation found π0 at roughly 72% success in-distribution, with only about a 5% drop under spatial perturbations and about 13% on unseen objects — strong by historical standards. Yet the dominant failure modes were pre-grasp, grasp and release/place errors, not failures to understand the command. And when a single pick instruction was replaced with a two-step command ("put the blue cube on top of the red cube"), first-step pick-up success fell from around 75% to about 35% — not because the language was hard, but because the physical precision and sequencing it implied were. [5]

Timing adds a second wall. Contact-rich tasks want control loops of tens of hertz so the robot can react to force changes; the largest VLAs run at single-digit hertz when the whole language model is in the loop, and getting a 50 Hz version requires architectural tricks such as flow-matching action heads rather than token-by-token decoding. [4][5]

So the honest update is: large models have made the "hard for humans" half *cheaper*, by giving robots commonsense-ish semantics almost for free. The "easy for humans" half — perception in the service of action, and dexterity under contact — remains the expensive part. The first half of Moravec's "easy for computers" claim has also been revised, since open-ended reasoning and common sense turned out not to be easy after all.

## How to read robot news differently

When you watch a robot demo, ask which half you're seeing.

If the robot understands an unusual command, finds an object in clutter it was never trained on, or reasons about which tool to use, you are watching the foundation-model side — genuine progress, but progress on the part that was previously *hard for humans*.

If the robot reliably folds a shirt, screws in a bolt, or hands over a full cup without spilling, you are watching the part Moravec flagged in 1988: contact, sensing, and the last centimeter.

Two further questions sharpen the picture. What is the success rate across new object positions, not just the best take? And how long does it take per attempt? A single success on video is a demonstration that a trajectory exists, not that a policy is robust.

Moravec's paradox ultimately says something simple: "intelligence" is not a single difficulty axis, and our intuitions about what is hard are calibrated to ourselves, not to the problem. Evolution spent something like a billion years making perception and movement look effortless, and that accumulated optimization is difficult to reverse-engineer. Where a machine finds a task easy depends on how small the search space is, how dense the feedback is, and how cheaply it can practice — and by those measures, chess was never the hard problem. Catching a ball was.

## Sources

1. [Wikipedia: Moravec's paradox — Moravec's 1988 quote, Minsky on unconscious competence, Pinker's summary, Brooks on early AI's definition of intelligence, and the evolutionary argument](https://en.wikipedia.org/wiki/Moravec%27s_paradox)
2. [Arvind Narayanan: Fact checking Moravec's paradox — sceptical review of the evolutionary story and of the claim that reasoning was ever easy for AI](https://www.normaltech.ai/p/fact-checking-moravecs-paradox)
3. [Understanding Moravec's Paradox — search space and reward sparsity as the machine-side account, with chess branching factor and game length figures](https://hexhowells.com/posts/moravecs-paradox.html)
4. [RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control — action-as-tokens recipe, generalization results, control frequency, and the stated limit that no new motions are acquired](https://robotics-transformer2.github.io/assets/rt2.pdf)
5. [Experiences from Benchmarking Vision–Language–Action Models for Robotic Manipulation — π0 success rates, dominant grasp/release failure modes, latency budgets, flow-matching action head](https://arxiv.org/html/2511.11298)
6. [A Survey on Imitation Learning for Contact-Rich Tasks in Robotics — nonlinear contact dynamics, sim-to-real gap widening with contact complexity, cost and safety limits of real-world RL](https://arxiv.org/html/2506.13498)
7. [Learning Dense Rewards for Contact-Rich Manipulation Tasks — sparse rewards yield no success within the training budget on peg-in-hole and USB insertion](https://ar5iv.labs.arxiv.org/html/2011.08458)
8. [OpenVLA: An Open-Source Vision-Language-Action Model — training data scale, multi-robot control, and typical control rates on new setups](https://openvla.github.io/)

---

Original article: https://eulore.ai/articles/moravecs-paradox-easy-things-hard-for-robots-529c3f32

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
