Tuesday 18 August 2026
botoïdes.com
FR·EN·DE·ES·IT·PT·NL·PL
The independent guide to home robots
Technical

VLA (Vision-Language-Action)

AI model that translates images and language into motor commands for a robot.

AI model architecture that merges visual understanding (images, video streams), natural language understanding and motor action generation into a single foundation neural network. A VLA takes camera images and verbal instructions as its input, and produces movement commands for a robot as its output. LG's CLOiD is one example: trained on tens of thousands of hours of household task data, it translates sight and speech into physical movements.

Articles about “VLA (Vision-Language-Action)”

Related terms

Join the discussion Soon

A question, a disagreement, hands-on feedback on this term? The Botoide community forum is coming soon.

← Full glossary