Fast head, slow backbone
mediummcq
Practice question
Why do several VLA systems run a large vision-language backbone at a few Hz while a small action head/decoder runs much faster?
Practice question
Why do several VLA systems run a large vision-language backbone at a few Hz while a small action head/decoder runs much faster?