Efficient Deep Learning Job Allocation in Cloud Systems by Predicting Resource Consumptions including GPU and CPU

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

One objective of GPU scheduling in cloud systems is to minimize the completion times of given deep learning models. This is important for deep learning in cloud environments because deep learning workloads require a lot of time to finish, and misallocation of these workloads can create a huge increase in job completion time. Difficulties of GPU scheduling come from a diverse type of parameters including model architectures and GPU types. Some of these model architectures are CPU-intensive rather than GPUintensive which creates a different hardware requirement when training different models. The previous GPU scheduling research had used a small set of parameters, which did not include CPU parameters, which made it difficult to reduce the job completion time (JCT). This paper introduces an improved GPU scheduling approach that reduces job completion time by predicting execution time and various resource consumption parameters including GPU Utilization%, GPU Memory Utilization%, GPU Memory, and CPU Utilization%. The experimental results show that the proposed model improves JCT by up to 40.9% on GPU Allocation based on Computing Efficiency compared to Driple.

키워드

cloud computing; convolutional neural network; deep learning; GPU job scheduling; performance estimation
제목
Efficient Deep Learning Job Allocation in Cloud Systems by Predicting Resource Consumptions including GPU and CPU
저자
Ferrino, Abuda Chad; Choe, Tae Young
DOI
10.31803/tg-20240112104444
발행일
2025-09
유형
Article
저널명
TEHNICKI GLASNIK-TECHNICAL JOURNAL
권
19
호
3
페이지
461 ~ 472