Why diffusion policies handle multimodal actions
mediummcq
Practice question
A demonstration set contains multiple valid ways to do a task (e.g. go left or right around an obstacle). Why does a diffusion (or other generative) policy typically handle this better than a network trained to regress actions with an MSE loss?