Mastering Atari with Discrete World Models

Hafner, Danijar; Lillicrap, Timothy P.; Norouzi, Mohammad; Ba, Jimmy

doi:10.48550/arxiv.2010.02193

Cited by 81 publications

(141 citation statements)

References 57 publications

(77 reference statements)

Supporting

Mentioning

139

Contrasting

Order By: Relevance

“…10, current SOTA algorithms like Agent57 may require more than 52.7 years of game-play to achieve SOTA performance, which revealed its low learning efficiency. As recommended in (Hafner et al 2020), we also argue for high learning efficiency algorithms, and we advocate that 200M training frames (equal to 38 days) are enough for achieving a superhuman agent.…”

Section: Current Challengesmentioning

confidence: 88%

“…In practice, we often use mean HNS or median HNS to show the final performance or generality of an algorithm. Dispute upon whether the mean value or the median value is more representative to show the generality and performance of the algorithms lasts for several years (Mnih et al 2015;Hessel et al 2017;Hafner et al 2020;Hessel et al 2021;Bellemare et al 2013;Machado et al 2018). To avoid any issues that aggregated metrics may have, we advocate calculating both of them in the final results because they serve different purposes, and we could not evaluate any algorithm via a single one of them.…”

Section: Normalized Scoresmentioning

confidence: 99%

“…Human World Records Baseline As (Toromanoff, Wirbel, and Moutarde 2019) put it, the Human Average Score Baseline potentially underestimates human performance relative to what is possible. To better reflect the performance of the algorithm compared to the human world record, we introduced a complete human world record baseline extended from (Hafner et al 2020;Toromanoff, Wirbel, and Moutarde 2019) to normalize the raw score, which is called the Human World Records Normalized Score (HWRNS), which can be calculated as follows:…”

Section: Normalized Scoresmentioning

confidence: 99%

“…Planning and Modeling Learning from sparse rewards is extremely difficult for model-free RL algorithms, especially those without intrinsic rewards that struggle to learn from weak gradient signals. Model-based methods can ease those problems by adopting a world model to planning (Schrittwieser et al 2020) or replay (Hafner et al 2020), which both enhanced the gradient signals. However, being utterly dependent on planning is unrealistic and will lose generality in some Atari Games like Tennis.…”

Section: Current Challengesmentioning

confidence: 99%

“…Perfection of the world records human baseline. We provide complete human world records overall the 57 Atari games, rather than part of them (Hafner et al 2020;Toromanoff, Wirbel, and Moutarde 2019). We further extended the SABER (Toromanoff, Wirbel, and Moutarde 2019) to a more comprehensive evaluation system with several new evaluation metrics based on human world records.…”

Section: Introductionmentioning

confidence: 99%

See 4 more Smart Citations

A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Fan¹

2021

Preprint

View full text Add to dashboard Cite

The Arcade Learning Environment (ALE) is proposed as an evaluation platform for empirically assessing the generality of agents across dozens of Atari 2600 games. ALE offers various challenging problems and has drawn significant attention from the deep reinforcement learning (RL) community. From Deep Q-Networks (DQN) to Agent57, RL agents seem to achieve superhuman performance in ALE. However, is this the case? In this paper, to explore this problem, we first review the current evaluation metrics in the Atari benchmarks and then reveal that the current evaluation criteria of achieving superhuman performance are inappropriate, which underestimated the human performance relative to what is possible. To handle those problems and promote the development of RL research, we propose a novel Atari benchmark based on human world records (HWR), which puts forward higher requirements for RL agents on both final performance and learning efficiency. Furthermore, we summarize the state-of-the-art (SOTA) methods in Atari benchmarks and provide benchmark results over new evaluation metrics based on human world records. We concluded that at least four open challenges hinder RL agents from achieving superhuman performance from those new benchmark results. Finally, we also discuss some promising ways to handle those problems.

show abstract

Section: Current Challengesmentioning

confidence: 88%

Section: Normalized Scoresmentioning

confidence: 99%

Section: Normalized Scoresmentioning

confidence: 99%

Section: Current Challengesmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

See 3 more Smart Citations

A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Fan¹

2021

Preprint

View full text Add to dashboard Cite

show abstract

Cycle-Consistent World Models for Domain Independent Latent Imagination

Bender

Joseph

Zöllner

2023

Lecture Notes in Computer Science

View full text Add to dashboard Cite

End-to-end autonomous driving seeks to solve the perception, decision, and control problems in an integrated way, which can be easier to generalize at scale and be more adapting to new scenarios. However, high costs and risks make it very hard to train autonomous cars in the real world. Simulations can therefore be a powerful tool to enable training. Due to slightly different observations, agents trained and evaluated solely in simulation often perform well there but have difficulties in real-world environments. To tackle this problem, we propose a novel model-based reinforcement learning approach called Cycleconsistent World Models. Contrary to related approaches, our model can embed two modalities in a shared latent space and thereby learn from samples in one modality (e.g., simulated data) and be used for inference in different domain (e.g., real-world data). Our experiments using different modalities in the CARLA simulator showed that this enables CCWM to outperform state-of-the-art domain adaptation approaches. Furthermore, we show that CCWM can decode a given latent representation into semantically coherent observations in both modalities.

show abstract

Inductive Biases in Machine Learning for Robotics and Control

Lutter

2023

Springer Tracts in Advanced Robotics

View full text Add to dashboard Cite

A fundamental problem of robotics is how can one program a robot to perform a task with its limited embodiment? Classical robotics solves this problem by carefully engineering interconnected modules. The main disadvantage is that this approach is labor-intensive and becomes close to impossible for unstructured environments and observations. Instead of manual engineering, one can solely use black-box models and data. In this paradigm, interconnected deep networks replace all modules of classical robotics. The network parameters are learned using reinforcement learning or self-supervised losses that predict the future.In this thesis, we want to show that these two approaches of classical engineering and black-box deep networks are not mutually exclusive. One can transfer insights from classical robotics to the black box deep networks and obtain better learning algorithms for robotics and control. To show that incorporating existing knowledge as inductive biases in machine learning algorithms can improve performance, we present three different algorithms: (1) The Differentiable Newton Euler Algorithm (Diff NEA) reinterprets the classical system identification of rigid bodies. By leveraging automatic differentiation, virtual parameters, and gradient-based optimization, this approach guarantees physically consistent parameters and applies to a wider class of dynamical systems. (2) Deep Lagrangian Networks (DeLaN) combines deep networks with Lagrangian mechanics to learn dynamics models that conserve energy. Using two networks to represent the potential and kinetic energy enables the computation of a physically plausible dynamics model using the Euler-Lagrange equation. (3) Robust Fitted Value Iteration (rFVI) leverages the control-affine dynamics of mechanical systems to extend value iteration to the adversarial reinforcement learning with continuous actions. The resulting approach enables the computation of the optimal policy that is robust to changes in the dynamics.Each of these algorithms is evaluated on physical systems and compared to the classical engineering and deep learning baselines. The experiments show that the inductive biases increase performance compared to black-box deep learning approaches. Diff NEA solves Ball-in-Cup on the physical Barrett WAM using offline model-based reinforcement learning and only four minutes of data. The deep networks models fail on this task despite using v vi• Jan Peters for being my supervisor. You cheered me up during the valleys, helped me celebrate the highs, always covered my back, increased my intrinsic motivation, and provided an excellent environment for me to complete my thesis. Without you, I could not have completed most of my goals for my thesis.• Russ Tedrake for agreeing to examine my thesis as well as the support of the other committee members, Kristian Kersting, Oskar van Stryk, and Stefan Roth.

show abstract

Mastering Atari with Discrete World Models

Cited by 81 publications

References 57 publications

A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

A Review for Deep Reinforcement Learning in Atari:Benchmarks, Challenges, and Solutions

Cycle-Consistent World Models for Domain Independent Latent Imagination

Inductive Biases in Machine Learning for Robotics and Control

Contact Info

Product

Resources

About