Skip to main content

Defining admissible rewards for high-confidence policy evaluation in batch reinforcement learning

Author(s): Prasad, Niranjani; Engelhardt, Barbara; Doshi-Velez, Finale

Download
To refer to this page use: http://arks.princeton.edu/ark:/88435/pr1zg13
Full metadata record
DC FieldValueLanguage
dc.contributor.authorPrasad, Niranjani-
dc.contributor.authorEngelhardt, Barbara-
dc.contributor.authorDoshi-Velez, Finale-
dc.date.accessioned2021-10-08T19:48:59Z-
dc.date.available2021-10-08T19:48:59Z-
dc.date.issued2020-04en_US
dc.identifier.citationPrasad, Niranjani, Barbara Engelhardt, and Finale Doshi-Velez. "Defining admissible rewards for high-confidence policy evaluation in batch reinforcement learning." In Proceedings of the ACM Conference on Health, Inference, and Learning (2020): pp. 1-9. doi:10.1145/3368555.3384450en_US
dc.identifier.urihttp://arks.princeton.edu/ark:/88435/pr1zg13-
dc.description.abstractA key impediment to reinforcement learning (RL) in real applications with limited, batch data is in defining a reward function that reflects what we implicitly know about reasonable behaviour for a task and allows for robust off-policy evaluation. In this work, we develop a method to identify an admissible set of reward functions for policies that (a) do not deviate too far in performance from prior behaviour, and (b) can be evaluated with high confidence, given only a collection of past trajectories. Together, these ensure that we avoid proposing unreasonable policies in high-risk settings. We demonstrate our approach to reward design on synthetic domains as well as in a critical care context, to guide the design of a reward function that consolidates clinical objectives to learn a policy for weaning patients from mechanical ventilation.en_US
dc.format.extent1 - 9en_US
dc.language.isoen_USen_US
dc.relation.ispartofProceedings of the ACM Conference on Health, Inference, and Learningen_US
dc.rightsFinal published version. This is an open access article.en_US
dc.titleDefining admissible rewards for high-confidence policy evaluation in batch reinforcement learningen_US
dc.typeConference Articleen_US
dc.identifier.doi10.1145/3368555.3384450-
pu.type.symplectichttp://www.symplectic.co.uk/publications/atom-terms/1.0/conference-proceedingen_US

Files in This Item:
File Description SizeFormat 
DefiningAdmissibleRewards.pdf1.18 MBAdobe PDFView/Download


Items in OAR@Princeton are protected by copyright, with all rights reserved, unless otherwise indicated.