Automatic Discovery and Transfer of Task Hierarchies in Reinforcement Learning

Mehta, Neville; Ray, Soumya; Tadepalli, Prasad; Dietterich, Thomas G.

doi:10.1609/aimag.v32i1.2342

Cited by 7 publications

(19 citation statements)

References 22 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Thus, s and correspond to different levels of a hierarchical situation model, and HRL provides methods to optimize option policies. While most HRL approaches assume a given pre-designed hierarchical structure [32, 36, 42, 57] or only bottom-up learning from the level of primitive states [53, 54, 58], our approach targets at general structural learning of behavioral and situation models by extending “is-a” and “has-parts” ontologies of situation models, including both specialization and generalization [16–18, 40]. …”

Section: Summary and Discussionmentioning

confidence: 99%

Toward Self-Referential Autonomous Learning of Object and Situation Models

et al. 2016

View full text Add to dashboard Cite

Most current approaches to scene understanding lack the capability to adapt object and situation models to behavioral needs not anticipated by the human system designer. Here, we give a detailed description of a system architecture for self-referential autonomous learning which enables the refinement of object and situation models during operation in order to optimize behavior. This includes structural learning of hierarchical models for situations and behaviors that is triggered by a mismatch between expected and actual action outcome. Besides proposing architectural concepts, we also describe a first implementation of our system within a simulated traffic scenario to demonstrate the feasibility of our approach.

show abstract

Section: Summary and Discussionmentioning

confidence: 99%

Toward Self-Referential Autonomous Learning of Object and Situation Models

et al. 2016

View full text Add to dashboard Cite

show abstract

“…Another promising approach that has been drawing much research interest is discovering the hierarchical structure automatically from state-action histories in the environment, either online or offline [Hengst 2002;Stolle 2004;Bakker and Schmidhuber 2004;Ş imşek et al 2005;Mehta et al 2008Mehta et al , 2011. For example, Mehta et al [2008] presents hierarchy induction via models and trajectories (HI-MAT), which discovers MAXQ task hierarchies by applying DBN models to successful execution trajectories of a source MDP task; the HEXQ [Hengst 2002[Hengst , 2004 method decomposes MDPs by finding nested sub-MDPs where there are policies to reach any exit with certainty; and Stolle [2004] performs automatic hierarchical decomposition by taking advantage of the factored representation of the underlying problem.…”

Section: Discussion: Maxq-op Algorithmmentioning

confidence: 99%

Online Planning for Large Markov Decision Processes with Hierarchical Decomposition

Bai

Chen

2015

ACM Trans. Intell. Syst. Technol.

View full text Add to dashboard Cite

Markov decision processes (MDPs) provide a rich framework for planning under uncertainty. However, exactly solving a large MDP is usually intractable due to the "curse of dimensionality"-the state space grows exponentially with the number of state variables. Online algorithms tackle this problem by avoiding computing a policy for the entire state space. On the other hand, since online algorithm has to find a near-optimal action online in almost real time, the computation time is often very limited. In the context of reinforcement learning, MAXQ is a value function decomposition method that exploits the underlying structure of the original MDP and decomposes it into a combination of smaller subproblems arranged over a task hierarchy. In this article, we present MAXQ-OP-a novel online planning algorithm for large MDPs that utilizes MAXQ hierarchical decomposition in online settings. Compared to traditional online planning algorithms, MAXQ-OP is able to reach much more deeper states in the search tree with relatively less computation time by exploiting MAXQ hierarchical decomposition online. We empirically evaluate our algorithm in the standard Taxi domain-a common benchmark for MDPs-to show the effectiveness of our approach. We have also conducted a long-term case study in a highly complex simulated soccer domain and developed a team named WrightEagle that has won five world champions and five runners-up in the recent 10 years of RoboCup Soccer Simulation 2D annual competitions. The results in the RoboCup domain confirm the scalability of MAXQ-OP to very large domains. ACM Reference Format:Aijun Bai, Feng Wu, and Xiaoping Chen. 2015. Online planning for large Markov decision processes with hierarchical decomposition.

show abstract

“…Some of the current HRL methods which are based on extracting the task-dependent hierarchy in FMDPs include HEX-Q [21], variable influence structure analysis (VISA) [22], and hierarchy induction via models and trajectories (HI-MAT) [23], [24]. Since there are implicit structure representations of the problems among the state variables in FMDPs, DBNs as a high-level source of pre-knowledge are often used to decompose the tasks in such processes, noting their capability to extract the impact of each action on the state variables.…”

Section: B Hrl Methods In Factored Mdps (Fmdps)mentioning

confidence: 99%

“…The state variables that affect others are assigned to deeper levels in the hierarchy. HI-MAT and VISA algorithms rely on the availability of DBNs for each action [22]- [24]. Since VISA considers the impacts of all actions regardless of the domain, it can create unnecessary branches in the extracted hierarchy or unnecessary sub-tasks.…”

Section: B Hrl Methods In Factored Mdps (Fmdps)mentioning

confidence: 99%

“…Since VISA considers the impacts of all actions regardless of the domain, it can create unnecessary branches in the extracted hierarchy or unnecessary sub-tasks. Thus, it may result in an ''exponentially sized hierarchy'' that limits its application in some domains [23], [24]. To address this problem, HI-MAT was proposed to remove such unsuccessful and redundant action cycles.…”

Section: B Hrl Methods In Factored Mdps (Fmdps)mentioning

confidence: 99%

See 1 more Smart Citation

Sequential Association Rule Mining for Autonomously Extracting Hierarchical Task Structures in Reinforcement Learning

2020

View full text Add to dashboard Cite

Reinforcement learning (RL) techniques, while often powerful, can suffer from slow learning speeds, particularly in high dimensional spaces or in environments with sparse rewards. The decomposition of tasks into a hierarchical structure holds the potential to significantly speed up learning, generalization, and transfer learning. However, the current task decomposition techniques often cannot extract hierarchical task structures without relying on high-level knowledge provided by an expert (e.g., using dynamic Bayesian networks (DBNs) in factored Markov decision processes), which is not necessarily available in autonomous systems. In this paper, we propose a novel method based on Sequential Association Rule Mining that can extract Hierarchical Structure of Tasks in Reinforcement Learning (SARM-HSTRL) in an autonomous manner for both Markov decision processes (MDPs) and factored MDPs. The proposed method leverages association rule mining to discover the causal and temporal relationships among states in different trajectories and extracts a task hierarchy that captures these relationships among sub-goals as termination conditions of different sub-tasks. We prove that the extracted hierarchical policy offers a hierarchically optimal policy in MDPs and factored MDPs. It should be noted that SARM-HSTRL extracts this hierarchical optimal policy without having dynamic Bayesian networks in scenarios with a single task trajectory and also with multiple tasks' trajectories. Furthermore, we show theoretically and empirically that the extracted hierarchical task structure is consistent with trajectories and provides the most efficient, reliable, and compact structure under appropriate assumptions. The numerical results compare the performance of the proposed SARM-HSTRL method with conventional HRL algorithms in terms of the accuracy in detecting the sub-goals, the validity of the extracted hierarchies, and the speed of learning in several testbeds. The key capabilities of SARM-HSTRL including handling multiple tasks and autonomous hierarchical task extraction can lead to the application of this HRL method in reusing, transferring, and generalization of knowledge in different domains. INDEX TERMS Association rule mining, extracting task structure, hierarchical reinforcement learning.

show abstract

Automatic Discovery and Transfer of Task Hierarchies in Reinforcement Learning

Cited by 7 publications

References 22 publications

Toward Self-Referential Autonomous Learning of Object and Situation Models

Toward Self-Referential Autonomous Learning of Object and Situation Models

Online Planning for Large Markov Decision Processes with Hierarchical Decomposition

Sequential Association Rule Mining for Autonomously Extracting Hierarchical Task Structures in Reinforcement Learning

Contact Info

Product

Resources

About