VLA (Vision-Language-Action)
AI model that translates images and language into motor commands for a robot.
AI model architecture that merges visual understanding (images, video streams), natural language understanding and motor action generation into a single foundation neural network. A VLA takes camera images and verbal instructions as its input, and produces movement commands for a robot as its output. LG's CLOiD is one example: trained on tens of thousands of hours of household task data, it translates sight and speech into physical movements.
Articles about “VLA (Vision-Language-Action)”
Related terms
Join the discussion Soon
A question, a disagreement, hands-on feedback on this term? The Botoide community forum is coming soon.
