Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A paper from McGill University introduces a novel class of off-policy algorithms, batch-constrained reinforcement learning, designed for learning from fixed data. The work demonstrates that standard algorithms like DQN and DDPG fail in this setting due to extrapolation errors. It presents the first continuous control deep reinforcement learning algorithm effective for arbitrary, fixed batch data.
License is listed as Open Access (green); specific terms should be verified.