How Humanoid Robots Learn

How humanoid robots actually learn to do tasks in 2026, from teleoperation and simulation to vision-language-action models, in plain terms.

Humanoid robots learn in three main ways in 2026, and most companies use all three together. Understanding them explains why a robot can fold a towel in a demo but still needs a human watching at home.

The first is teleoperation. A person remotely controls the robot through a task, and the robot records what good performance looks like. This is how 1X’s NEO handles chores it has not mastered, and how Sanctuary trains Phoenix. Every guided run is a training example, which is why some home robots ship before they are fully autonomous. The human is both the helper and the teacher.

The second is simulation. Companies run robots through millions of virtual repetitions in a physics engine, far faster and cheaper than real life, then transfer what works to the physical robot. The gap between simulation and reality, called sim-to-real, is one of the field’s central engineering problems, and Chinese firms in particular have invested heavily in closing it.

The third is the vision-language-action model, the approach behind Figure’s Helix and similar systems. The robot takes in what it sees, understands a goal stated in plain language, and outputs movement directly, learning general skills rather than scripted routines.

The reason none of this is solved: there is no internet-sized archive of physical tasks the way there is for text. Data is scarce and expensive to gather, so robots learn slowly and unevenly. Progress in humanoids is really progress in data and learning, more than in motors and frames.