Reward shaping without changing the optimal policy
hardsubjective
Practice question
You want to add shaping rewards to speed up RL training. How can you add them without changing the optimal policy, and what goes wrong with naive shaping?
Voice input needs Chrome or Edge. On this browser, type your answer — grading is the same.