Reinforcement learning

Instead of a programmer writing rules, the engineer writes a reward function, for example plus points for moving forward at the target speed, minus points for falling, for wasting energy or for jerky joints. An algorithm then explores actions and adjusts a neural network policy towards behaviour that scores higher.

Real robots break, so most training happens in physics simulation where thousands of copies can run faster than real time. The learned policy is then transferred to hardware, a step known as sim-to-real.

Quadruped and humanoid locomotion controllers are the clearest success: a policy trained on randomised terrain recovers from slips and shoves in ways hand-written rules rarely manage. Reward design is the hard part, since a badly specified reward produces technically optimal but useless behaviour.

Your premier source for robotics news, AI innovations, and automation technology insights.

© 2026 RoboterGalaxy. All rights reserved.