Robotics

Reward shaping without changing the optimal policy

hardsubjective

Practice question

You want to add shaping rewards to speed up RL training. How can you add them without changing the optimal policy, and what goes wrong with naive shaping?

Voice input needs Chrome or Edge. On this browser, type your answer — grading is the same.