Vision-language-action model

VLAs extend the recipe behind large language and vision models to control. They are trained on large sets of robot demonstrations, images paired with the commands that were executed, often together with ordinary web image and text data so the model inherits general knowledge about objects.

At runtime the model receives the current camera view and a text instruction and predicts the next action, typically a gripper pose or joint change, many times per second. The appeal is generality: instead of a separate program per task, one model can attempt tasks phrased in words, including some it has not seen in exactly that form.

Published examples include Google DeepMind RT-2 and the open-source OpenVLA. Reliability, speed and the need for large, diverse robot datasets remain the open problems.

Your premier source for robotics news, AI innovations, and automation technology insights.

© 2026 RoboterGalaxy. All rights reserved.