Green Resource Allocation Based on Deep Reinforcement Learning in Content-Centric IoT

He, Xiaoming; Wang, Kun; Huang, Huawei; Miyazaki, Toshiaki; Wang, Yixuan; Guo, Song

doi:10.1109/tetc.2018.2805718

Cited by 182 publications

(85 citation statements)

References 56 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Secondly, the dueling DQN approach is also integrated in the design with the intuition that it is not always necessary to estimate the reward by taking some action. The state-action Qvalue in dueling DQN is decomposed into one value function representing the reward in the current state, and the advantage optimize cache hit rate [70], cache expiration time [74], interference alignment [76]- [78], Quality of Experience [79], [81], energy efficiency [84], resource allocation [85]- [87], traffic latency, or redundancy [89], [91]. function that measures the relative importance of a certain action compared with other actions.…”

Section: A Wireless Proactive Cachingmentioning

confidence: 99%

“…The TE-aware exploration leverages the shortest path algorithm and NUM-based solution as the baseline during exploration. The PER method is conventionally used in DQL, e.g., [79] and [109], while the authors in [136] integrate the PER method with the actor-critic framework for the first time. The proposed scheme assigns different priorities to transitions in the experience replay.…”

Section: A Traffic Engineering and Routingmentioning

confidence: 99%

See 1 more Smart Citation

Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Luong

Hoang

Gong

et al. 2019

IEEE Commun. Surv. Tutorials

1,305

475

View full text Add to dashboard Cite

This paper presents a comprehensive literature review on applications of deep reinforcement learning in communications and networking. Modern networks, e.g., Internet of Things (IoT) and Unmanned Aerial Vehicle (UAV) networks, become more decentralized and autonomous. In such networks, network entities need to make decisions locally to maximize the network performance under uncertainty of network environment. Reinforcement learning has been efficiently used to enable the network entities to obtain the optimal policy including, e.g., decisions or actions, given their states when the state and action spaces are small. However, in complex and large-scale networks, the state and action spaces are usually large, and the reinforcement learning may not be able to find the optimal policy in reasonable time. Therefore, deep reinforcement learning, a combination of reinforcement learning with deep learning, has been developed to overcome the shortcomings. In this survey, we first give a tutorial of deep reinforcement learning from fundamental concepts to advanced models. Then, we review deep reinforcement learning approaches proposed to address emerging issues in communications and networking. The issues include dynamic network access, data rate control, wireless caching, data offloading, network security, and connectivity preservation which are all important to next generation networks such as 5G and beyond. Furthermore, we present applications of deep reinforcement learning for traffic routing, resource sharing, and data collection. Finally, we highlight important challenges, open issues, and future research directions of applying deep reinforcement learning.

show abstract

Section: A Wireless Proactive Cachingmentioning

confidence: 99%

Section: A Traffic Engineering and Routingmentioning

confidence: 99%

Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Luong

Hoang

Gong

et al. 2019

IEEE Commun. Surv. Tutorials

1,305

475

View full text Add to dashboard Cite

show abstract

“…for each decision epoch k = 0, 1, 2, · · · do k ← k + 1 for each action a k ∈ A s k do Determine the post-decision global system states k ← f (s k , a k ) by (5) Determine the post-decision local system state vector {s n,k } N n=1 ← {f n (s n,k , a k )} N n=1 by (15) for each IoT device n ∈ N do Encodes k in the input vector xs n,k Determine the feature φs n,k (s k ) by (20), (21), (22) Determine the local reward functiong n (s k ) by (17) end for Determine the value of the RHS of (25) end for Select the action according to (25) end for…”

Section: Algorithm 1 Determine the Optimal Actionmentioning

confidence: 99%

“…The detailed procedures and equations to update the value functions and weights are given in Appendix B and summarized in Algorithm 2 below. Note that at the kth decision epoch, θ n,k is used in place of θ n in (25) to derive the optimal action π * (s k ). In the following discussion, we add a subscript k to the notations described in Section III.B to represent the parameter values at the k-th decision epoch.…”

Section: Per-node Value Function and Weight Updatementioning

confidence: 99%

“…• Neural-ICO algorithm: It is similar to the DQN algorithm used in [23], [25], except that the output from the neural network is the value functions for the post-decision states instead of the Q factors for the state and action pairs. The function approximation architecture of this algorithm is given in Fig.11 in Appendix C. We developed a discrete-event system-level simulator for the NB-IoT MEC system, where the the simulation parameters are given in Table II.…”

Section: Remark 2 (Signaling Overhead Of Semi-distributed Implementatmentioning

confidence: 99%

See 1 more Smart Citation

Multiuser Resource Control With Deep Reinforcement Learning in IoT Edge Computing

Lei

Xiong

et al. 2019

IEEE Internet Things J.

View full text Add to dashboard Cite

By leveraging the concept of mobile edge computing (MEC), massive amount of data generated by a large number of Internet of Things (IoT) devices could be offloaded to MEC server at the edge of wireless network for further computational intensive processing. However, due to the resource constraint of IoT devices and wireless network, both the communications and computation resources need to be allocated and scheduled efficiently for better system performance. In this paper, we propose a joint computation offloading and multi-user scheduling algorithm for IoT edge computing system to minimize the long-term average weighted sum of delay and power consumption under stochastic traffic arrival. We formulate the dynamic optimization problem as an infinite-horizon average-reward continuous-time Markov decision process (CTMDP) model. One critical challenge in solving this MDP problem for the multi-user resource control is the curse-of-dimensionality problem, where the state space of the MDP model and the computation complexity increase exponentially with the growing number of users or IoT devices. In order to overcome this challenge, we use the deep reinforcement learning (RL) techniques and propose a neural network architecture to approximate the value functions for the post-decision system states. The designed algorithm to solve the CTMDP problem supports semi-distributed auction-based implementation, where the IoT devices submit bids to the BS to make the resource control decisions centrally. Simulation results show that the proposed algorithm provides significant performance improvement over the baseline algorithms, and also outperforms the RL algorithms based on other neural network architectures.

show abstract

DRL at the Application and Service Layer

2023

Deep Reinforcement Learning for Wireless Communications and Networking

View full text Add to dashboard Cite

Green Resource Allocation Based on Deep Reinforcement Learning in Content-Centric IoT

Cited by 182 publications

References 56 publications

Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Applications of Deep Reinforcement Learning in Communications and Networking: A Survey

Multiuser Resource Control With Deep Reinforcement Learning in IoT Edge Computing

DRL at the Application and Service Layer

Contact Info

Product

Resources

About