RL without TD Learning: A New 'Divide and Conquer' Algorithm
Researchers from BAIR (Berkeley AI) have introduced the Transitive RL (TRL) algorithm, based on a 'divide and conquer' paradigm for off-policy reinforcement learning (RL). TRL reduces the number of Bellman recursions logarithmically, avoiding the error accumulation problems typical of TD learning. The algorithm showed superior results on complex long-horizon tasks without requiring tuning of the hyperparameter n as in n-step TD.
Berkeley Artificial Intelligence Research (BAIR)
BAIR (Berkeley AI)27.07 · 17:04
