Robotics

Action chunking in VLAs

mediumsubjective

Practice question

Many vision-language-action models predict a short sequence (chunk) of future actions per inference rather than a single action. Why?

Voice input needs Chrome or Edge. On this browser, type your answer — grading is the same.