Jellyfish: Timely Inference Serving for Dynamic Edge Networks

Nigade, Vinod; Bauszat, Pablo; Bal, Henri E.; Wang, Lin

doi:10.1109/rtss55097.2022.00032

Cited by 10 publications

(3 citation statements)

References 31 publications

(42 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…ML inference services are user-facing, which mandates high responsiveness [18,37]. Moreover, high accuracy is crucial for these services [20,26]. Consequently, inference systems must deliver highly accurate predictions with fewer computing resources (cost-efficient) while meeting latency constraints under workload variations [18,19,29,37].…”

Section: Featurementioning

confidence: 99%

“…Conversely, overprovisioning wastes computing resources [30,36]. To address these problems caused by dynamic workloads, Autoscaling [2,9,17,18,30,36] resizes the resources of the service, and Model-switching [26,38] switches between ML model variants that differ in their inference latency and accuracy (higher accuracy, higher latency); the former tries to be cost-efficient, and the latter tries to be more accurate, while both guarantee latency SLOs.…”

Section: Featurementioning

confidence: 99%

“…Due to interaction with online users, inference services are latency-sensitive [18,36], and since they contain heavy computations, they are resource-intensive [10]. Accuracy is also a pillar dimension of these services [26]. Faced with dynamic workload [18,37], it is essential to consider the ternary tradeoff space between latency, accuracy, and the resource cost dynamically to address latency requirements while gaining higher accuracy cost-efficiently.…”

Section: Motivationmentioning

confidence: 99%

See 2 more Smart Citations

Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems

Salmani

Ghafouri

Sanaee

et al. 2023

Proceedings of the 3rd Workshop on Machine Learning and Systems

View full text Add to dashboard Cite

The use of machine learning (ML) inference for various applications is growing drastically. ML inference services engage with users directly, requiring fast and accurate responses. Moreover, these services face dynamic workloads of requests, imposing changes in their computing resources. Failing to right-size computing resources results in either latency service level objectives (SLOs) violations or wasted computing resources. Adapting to dynamic workloads considering all the pillars of accuracy, latency, and resource cost is challenging. In response to these challenges, we propose InfAdapter, which proactively selects a set of ML model variants with their resource allocations to meet latency SLO while maximizing an objective function composed of accuracy and cost. InfAdapter decreases SLO violation and costs up to 65% and 33%, respectively, compared to a popular industry autoscaler (Kubernetes Vertical Pod Autoscaler).

show abstract