Multi-Game Decision Transformers

Lee, Kuang-Huei; Nachum, Ofir; Yang, Ming‐Bo; Lee, Lisa; Freeman, Daniel; Xu, Winnie; Guadarrama, Sergio; Fischer, Ian; Michalewski, Henryk; Mordatch, Igor

doi:10.48550/arxiv.2205.15241

Cited by 7 publications

(28 citation statements)

References 34 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…where t is a time step and R is the return for the remaining sequence. The sequence we consider here is similar to the one used in [30] whereas we do not include reward as part of the sequence and we predict an additional quantity R that enables us to estimate an optimal input length, which we will cover in the following paragraphs. Figure 2 presents an overview of our model architecture.…”

Section: Reinforcement Learning As Sequence Modelingmentioning

confidence: 99%

“…Our training method extends the work of [30] by estimating the maximum expected return value for a trajectory using Equation 2. This estimation aids in comparing expected returns of different trajectories over various history lengths.…”

Section: Training Objective For Maximum In-support Returnmentioning

confidence: 99%

“…To sample from expert return distribution P (R t , ...|expert t ), we adopt an approach similar to [30] by applying Bayes' rule P (R t , ...|expert t ) ∝ P (expert t |R t , ...)P (R t , ...) and approximate the distribution of expert-level return with inverse temperature κ 1 [24,48,47,43]:…”

Section: Action Inference During Test Timementioning

confidence: 99%

“…To ensure fair comparisons, all methods employ the same architecture for the image encoder. Following [30], we incorporate random cropping and rotation for image augmentation. Additional experiment details are delegated to the Appendix for brevity.…”

Section: Multi-task Offline Reinforcement Learningmentioning

confidence: 99%

See 3 more Smart Citations

Multi-focus image fusion: Transformer and shallow feature attention matters

Jiang

Hua

et al. 2023

Displays

View full text Add to dashboard Cite

Section: Reinforcement Learning As Sequence Modelingmentioning

confidence: 99%

Section: Training Objective For Maximum In-support Returnmentioning

confidence: 99%

Section: Action Inference During Test Timementioning

confidence: 99%

Section: Multi-task Offline Reinforcement Learningmentioning

confidence: 99%

See 2 more Smart Citations

Multi-focus image fusion: Transformer and shallow feature attention matters

Jiang

Hua

et al. 2023

Displays

View full text Add to dashboard Cite

“…When learning a predictive information representation between a state, action pair and its subsequent state, the learning task is equivalent to modeling environment dynamics [4]. In this work, we are interested in training multi-task generalist agents [5,6,7] that can master a wide range of robotics skills in both simulated and real environments by learning from a large amount of diverse experience. We hypothesize that modeling the predictive information will give latent representations that capture environment dynamics across multiple tasks, making it simpler and more efficient to learn a generalist policy.…”

Section: Introductionmentioning

confidence: 99%

PI-QT-Opt: Predictive Information Improves Multi-Task Robotic Reinforcement Learning at Scale

Lee¹,

Ted²,

Li³

et al. 2022

Preprint

View full text Add to dashboard Cite

The predictive information, the mutual information between the past and future, has been shown to be a useful representation learning auxiliary loss for training reinforcement learning agents, as the ability to model what will happen next is critical to success on many control tasks. While existing studies are largely restricted to training specialist agents on single-task settings in simulation, in this work, we study modeling the predictive information for robotic agents and its importance for general-purpose agents that are trained to master a large repertoire of diverse skills from large amounts of data. Specifically, we introduce Predictive Information QT-Opt (PI-QT-Opt), a QT-Opt agent augmented with an auxiliary loss that learns representations of the predictive information to solve up to 297 vision-based robot manipulation tasks in simulation and the real world with a single set of parameters. We demonstrate that modeling the predictive information significantly improves success rates on the training tasks and leads to better zeroshot transfer to unseen novel tasks. Finally, we evaluate PI-QT-Opt on real robots, achieving substantial and consistent improvement over QT-Opt in multiple experimental settings of varying environments, skills, and multi-task configurations.

show abstract