Swayam

Gujarati, Arpan; Elnikety, Sameh; He, Yuxiong; McKinley, Kathryn S.; Brandenburg, Björn B.

doi:10.1145/3135974.3135993

Cited by 75 publications

(23 citation statements)

References 18 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…MS [38] INFaaS [30] Cocktail [20] VPA [9] InfAdapter Cost Optimization ✕ ✓ ✓ * ✓ ✓ Accuracy Maximization ✓ ✕ ✓ ✕ ✓ Predictive Decision-Making ✓ ✕ ✓ ✓ ✓ Container as a Service (CaaS) ✕ ✕ ✕ ✓ ✓ Latency SLO-aware ✓ ✓ ✓ ✕ ✓ machine translation, chatbots, medical, and recommender systems, are running in data centers [13,28,32,34], comprising more than 90% of computing resources allocated to ML [10,13,25]. ML inference services are user-facing, which mandates high responsiveness [18,37]. Moreover, high accuracy is crucial for these services [20,26].…”

Section: Featurementioning

confidence: 99%

“…Moreover, high accuracy is crucial for these services [20,26]. Consequently, inference systems must deliver highly accurate predictions with fewer computing resources (cost-efficient) while meeting latency constraints under workload variations [18,19,29,37].…”

Section: Featurementioning

confidence: 99%

“…The dynamic nature of inference serving workloads requires different resource allocations for ML services [18,37]. Failing to right-size the services results in over or underresource provisioning.…”

Section: Featurementioning

confidence: 99%

“…Failing to right-size the services results in over or underresource provisioning. Under-provisioning leads to service level objective (SLO) violations (e.g., 99 𝑡ℎ percentile of latency distribution, P99-latency) [18,36]. Conversely, overprovisioning wastes computing resources [30,36].…”

Section: Featurementioning

confidence: 99%

“…Conversely, overprovisioning wastes computing resources [30,36]. To address these problems caused by dynamic workloads, Autoscaling [2,9,17,18,30,36] resizes the resources of the service, and Model-switching [26,38] switches between ML model variants that differ in their inference latency and accuracy (higher accuracy, higher latency); the former tries to be cost-efficient, and the latter tries to be more accurate, while both guarantee latency SLOs.…”

Section: Featurementioning

confidence: 99%

See 4 more Smart Citations

Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems

Salmani

Ghafouri

Sanaee

et al. 2023

Proceedings of the 3rd Workshop on Machine Learning and Systems

View full text Add to dashboard Cite

The use of machine learning (ML) inference for various applications is growing drastically. ML inference services engage with users directly, requiring fast and accurate responses. Moreover, these services face dynamic workloads of requests, imposing changes in their computing resources. Failing to right-size computing resources results in either latency service level objectives (SLOs) violations or wasted computing resources. Adapting to dynamic workloads considering all the pillars of accuracy, latency, and resource cost is challenging. In response to these challenges, we propose InfAdapter, which proactively selects a set of ML model variants with their resource allocations to meet latency SLO while maximizing an objective function composed of accuracy and cost. InfAdapter decreases SLO violation and costs up to 65% and 33%, respectively, compared to a popular industry autoscaler (Kubernetes Vertical Pod Autoscaler).

show abstract

Section: Featurementioning

confidence: 99%

Section: Featurementioning

confidence: 99%

Section: Featurementioning

confidence: 99%

Section: Featurementioning

confidence: 99%

Section: Featurementioning

confidence: 99%

See 3 more Smart Citations

Reconciling High Accuracy, Cost-Efficiency, and Low Latency of Inference Serving Systems

Salmani

Ghafouri

Sanaee

et al. 2023

Proceedings of the 3rd Workshop on Machine Learning and Systems

View full text Add to dashboard Cite

show abstract

ProKube: Proactive Kubernetes Orchestrator for Inference in Heterogeneous Edge Computing

Ali,

Golec,

Singh Gill

et al. 2024

Int J Network Mgmt

View full text Add to dashboard Cite

Deep neural network (DNN) and machine learning (ML) models/ inferences produce highly accurate results demanding enormous computational resources. The limited capacity of end‐user smart gadgets drives companies to exploit computational resources in an edge‐to‐cloud continuum and host applications at user‐facing locations with users requiring fast responses. Kubernetes hosted inferences with poor resource request estimation results in service level agreement (SLA) violation in terms of latency and below par performance with higher end‐to‐end (E2E) delays. Lifetime static resource provisioning either hurts user experience for under‐resource provisioning or incurs cost with over‐provisioning. Dynamic scaling offers to remedy delay by upscaling leading to additional cost whereas a simple migration to another location offering latency in SLA bounds can reduce delay and minimize cost. To address this cost and delay challenges for ML inferences in the inherent heterogeneous, resource‐constrained, and distributed edge environment, we propose ProKube, which is a proactive container scaling and migration orchestrator to dynamically adjust the resources and container locations with a fair balance between cost and delay. ProKube is developed in conjunction with Google Kubernetes Engine (GKE) enabling cross‐cluster migration and/ or dynamic scaling. It further supports the regular addition of freshly collected logs into scheduling decisions to handle unpredictable network behavior. Experiments conducted in heterogeneous edge settings show the efficacy of ProKube to its counterparts cost greedy (CG), latency greedy (LG), and GeKube (GK). ProKube offers 68%, 7%, and 64% SLA violation reduction to CG, LG, and GK, respectively, and it improves cost by 4.77 cores to LG and offers more cost of 3.94 to CG and GK.

show abstract