Offline Reinforcement Learning: Addressing Out-of-Distribution Actions


Dr. Swakkhar Shatabda (SWK)

Professor

swakkhar.shatabda@bracu.ac.bd

Synopsis

 

A major challenge in offline reinforcement learning (RL) is: how to learn an effective policy using only a previously collected dataset, without further interaction with the environment. Since the agent cannot explore and collect new data, it may generate out-of-distribution (OOD) actions, actions that are poorly represented or completely absent in the training dataset. Such actions can lead to unreliable or overly optimistic decisions.  In general, the problem can be summarized as:

 

How can an offline RL agent generate high-reward actions while remaining within the distribution of actions supported by the available dataset?

 


Relevant courses to the topic

 

  • Probability and Statistics, Reinforcement Learning

 


Reading List

 

  • Hu, X., Li, S., Xu, Y., Tang, B., & Chen, L. (2026). Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 40(26), 21894–21902. https://doi.org/10.1609/aaai.v40i26.39342

 



©2026 BracU CSE Department