Skip to main navigation Skip to search Skip to main content

VOQL: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation

  • Alekh Agarwal
  • , Yujia Jin
  • , Tong Zhang

Research output: Chapter in Book/Conference Proceeding/ReportConference Paper published in a bookpeer-review

Abstract

We study time-inhomogeneous episodic reinforcement learning (RL) under general function approximation and sparse rewards. We design a new algorithm, Variance-weighted Optimistic QLearning (VOQL), based on Q-learning and bound its regret assuming closure under Bellman backups, and bounded Eluder dimension for the regression function class. As a special case, VOQL achieves Oe(d√TH + d6H5) regret over T episodes for a horizon H MDP under (ddimensional) linear function approximation, which is asymptotically optimal. Our algorithm incorporates weighted regression-based upper and lower bounds on the optimal value function to obtain this improved regret. The algorithm is computationally efficient given a regression oracle over the function class, making this the first computationally tractable and statistically optimal approach for linear MDPs.

Original languageEnglish
Title of host publicationProceedings of Thirty Sixth Conference on Learning Theory
PublisherML Research Press
Pages987-1063
Number of pages77
Volume195
Publication statusPublished - 12 Jul 2023
Externally publishedYes
Event36th Annual Conference on Learning Theory, COLT 2023 - Bangalore, India
Duration: 12 Jul 202315 Jul 2023

Publication series

NameProceedings of Machine Learning Research
PublisherML Research Press
Volume195
ISSN (Print)2640-3498

Conference

Conference36th Annual Conference on Learning Theory, COLT 2023
Country/TerritoryIndia
CityBangalore
Period12/07/2315/07/23

Bibliographical note

Publisher Copyright:
© 2023 A. Agarwal, Y. Jin & T. Zhang.

Keywords

  • Reinforcement learning
  • nonlinear function approximation
  • model-free algorithms
  • eluder dimension

Fingerprint

Dive into the research topics of 'VOQL: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation'. Together they form a unique fingerprint.

Cite this