Off-Policy Evaluation
离线策略评估OPEAdvancedEstimating how much return a new policy would actually get once deployed, using only data collected by other policies.
Off-policy evaluation is a class of problem in reinforcement learning: data was collected by some behavior policy, and the goal is to estimate the expected return of a different target policy without letting it interact with the environment. Testing a policy on a real robot requires a human minder and wears down hardware, so if bad policies can be screened out using existing logs first, it saves a great deal of evaluation cost, and it also helps offline reinforcement learning pick checkpoints and hyperparameters. Common methods include importance sampling (reweighting old data by the ratio of the probability that each policy would pick the same action), fitted Q evaluation (FQE, which fits the target policy's Q-function from data), doubly robust estimation, and model-based approaches. Fu and colleagues' 2021 DOPE benchmark judges methods not just by value-estimation error but also by rank correlation and regret@k, since in practice getting the ranking right usually matters more than getting the exact number right.
ExampleGoogle's Irpan and colleagues (NeurIPS 2019) reframed OPE as a classification problem for image-based robotic grasping, and using only offline data were able to reliably predict the relative performance of several policies on a real robot, including in sim-to-real transfer settings.
- Also called
- OPE, Off-Policy Policy Evaluation
- Related
- Offline Reinforcement Learning · Off-Policy · Q-Function · Real-World Evaluation · World-Model-based Policy Evaluation · Importance Sampling
- Sources
- Benchmarks for Deep Off-Policy Evaluation (DOPE, ICLR 2021)
Off-Policy Evaluation via Off-Policy Classification (NeurIPS 2019)