On-policy vs off-policy RL
mediumsubjective
Practice question
Define on-policy and off-policy RL, give an algorithm example of each, and explain the sample-efficiency trade-off that matters for robot learning.
Voice input needs Chrome or Edge. On this browser, type your answer — grading is the same.