Balancing information with observation costs in deep reinforcement learning

From National Research Council Canada

Download	View final version: Balancing information with observation costs in deep reinforcement learning (PDF, 2.2 MiB)
Link	https://caiac.pubpub.org/pub/0jmy7gpd/release/1
Author	Search for: Bellinger, Colin¹; Search for: Drozdyuk, Andriy¹; Search for: Crowley, Mark; Search for: Tamblyn, Isaac²
Affiliation	National Research Council of Canada. Digital Technologies National Research Council of Canada. Security and Disruptive Technologies
Format	Text, Article
Conference	The 35th Canadian Conference on Artificial Intelligence, May 30 - June 3, 2022, Toronto, ON., Virtual
Physical description	12 p.
Subject	deep reinforcement learning; partial observability; state measurement costs
Abstract	The use of reinforcement learning (RL) in scientific applications, such as materials design and automated chemistry, is increasing. A major challenge, however, lies in fact that measuring the state of the system is often costly and time consuming in scientific applications, whereas policy learning with RL requires a measurement after each time step. In this work, we make the measurement costs explicit in the form of a costed reward and propose the active-measure with costs framework that enables off-the-shelf deep RL algorithms to learn a policy for both selecting actions and determining whether or not to measure the state of the system at each time step. In this way, the agents learn to balance the need for information with the cost of information. Our results show that when trained under this regime, the Dueling DQN and PPO agents can learn optimal action policies whilst making up to 50\% fewer state measurements, and recurrent neural networks can produce a greater than 50\% reduction in measurements. We postulate the these reduction can help to lower the barrier to applying RL to real-world scientific applications.
Publication date	2022-05-27
Publisher	Canadian Artificial Intelligence Association
Licence	Creative Commons, Attribution 4.0 International (CC BY 4.0) https://creativecommons.org/licenses/by/4.0/
In	Proceedings of the 35th Canadian Conference on Artificial Intelligence, 2022L5 (27 May 2022). https://caiac.pubpub.org/ai2022.
Language	English
Peer reviewed	Yes
Export citation	Export as RIS
Report a correction	Report a correction (opens in a new tab)
Record identifier	2e701a0c-744b-4c24-b5ce-b8191023ca33
Record created	2022-06-22
Record modified	2022-06-22

Date modified:: 2024-07-17