Pushi Zhang scite author profile

Pushi Zhang

4Publications

5Citation Statements Received

15Citation Statements Given

How they've been cited

How they cite others

Affiliations

University of Pennsylvania, Tsinghua University

Publications

Order By: Most citations

Independence-aware Advantage Estimation

Zhang

Liu

et al. 2021

View full text Add to dashboard Cite

Most of existing advantage function estimation methods in reinforcement learning suffer from the problem of high variance, which scales unfavorably with the time horizon. To address this challenge, we propose to identify the independence property between current action and future states in environments, which can be further leveraged to effectively reduce the variance of the advantage estimation. In particular, the recognized independence property can be naturally utilized to construct a novel importance sampling advantage estimator with close-to-zero variance even when the Monte-Carlo return signal yields a large variance. To further remove the risk of the high variance introduced by the new estimator, we combine it with existing Monte-Carlo estimator via a reward decomposition model learned by minimizing the estimation variance. Experiments demonstrate that our method achieves higher sample efficiency compared with existing advantage estimation methods in complex environments.

show abstract

Distributional Reinforcement Learning for Multi-Dimensional Reward Functions

Zhang¹,

Chen²,

Li³

et al. 2021

Preprint

View full text Add to dashboard Cite

A growing trend for value-based reinforcement learning (RL) algorithms is to capture more information than scalar value functions in the value network. One of the most well-known methods in this branch is distributional RL, which models return distribution instead of scalar value. In another line of work, hybrid reward architectures (HRA) in RL have studied to model source-specific value functions for each source of reward, which is also shown to be beneficial in performance. To fully inherit the benefits of distributional RL and hybrid reward architectures, we introduce Multi-Dimensional Distributional DQN (MD3QN), which extends distributional RL to model the joint return distribution from multiple reward sources. As a by-product of joint distribution modeling, MD3QN can capture not only the randomness in returns for each source of reward, but also the rich reward correlation between the randomness of different sources. We prove the convergence for the joint distributional Bellman operator and build our empirical algorithm by minimizing the Maximum Mean Discrepancy between joint return distribution and its Bellman target. In experiments, our method accurately models the joint return distribution in environments with richly correlated reward functions, and outperforms previous RL methods utilizing multi-dimensional reward functions in the control setting.

show abstract

Demonstration actor critic

Liu

Zhang

et al. 2021

Neurocomputing

View full text Add to dashboard Cite

An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable Context

Chen¹,

Zhu²,

Zheng³

et al. 2022

Preprint

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Pushi Zhang

Independence-aware Advantage Estimation

Distributional Reinforcement Learning for Multi-Dimensional Reward Functions

Demonstration actor critic

An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable Context

Contact Info

Product

Resources

About