DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters

PENG, Y; BAO, Y; CHEN, Y; Wu, C; Meng, C; Lin, W

File Download

re01.htm

Links for fulltext

(May Require Subscription)

Publisher Website: 10.1109/TPDS.2021.3052895
Scopus: eid_2-s2.0-85099732170
WOS: WOS:000622094200004
Find via

Supplementary

Citations:
- Scopus: 0
- Web of Science: 0
Appears in Collections:
- Computer Science: Journal/Magazine Articles

Article: DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters

Title	DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
Authors	PENG, Y BAO, Y CHEN, Y Wu, C Meng, C Lin, W
Keywords	Deep learning resource allocation distributed training
Issue Date	2021
Publisher	Institute of Electrical and Electronics Engineers. The Journal's web site is located at http://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=71
Citation	IEEE Transactions on Parallel and Distributed Systems, 2021, v. 32 n. 8, p. 1947-1960 How to Cite? DOI: http://dx.doi.org/10.1109/TPDS.2021.3052895
Abstract	Efficient resource scheduling is essential for maximal utilization of expensive deep learning (DL) clusters. Existing cluster schedulers either are agnostic to machine learning (ML) workload characteristics, or use scheduling heuristics based on operators' understanding of particular ML framework and workload, which are less efficient or not general enough. In this article, we show that DL techniques can be adopted to design a generic and efficient scheduler. Specifically, we propose DL2, a DL-driven scheduler for DL clusters, targeting global training job expedition by dynamically resizing resources allocated to jobs. DL2 advocates a joint supervised learning and reinforcement learning approach: a neural network is warmed up via offline supervised learning based on job traces produced by the existing cluster scheduler; then the neural network is plugged into the live DL cluster, fine-tuned by reinforcement learning carried out throughout the training progress of the DL jobs, and used for deciding job resource allocation in an online fashion. We implement DL2 on Kubernetes and enable dynamic resource scaling in DL jobs on MXNet. Extensive evaluation shows that DL2 outperforms fairness scheduler (i.e., DRF) by 44.1 percent and expert heuristic scheduler (i.e., Optimus) by 17.5 percent in terms of average job completion time.
Persistent Identifier	http://hdl.handle.net/10722/301454
ISSN	1045-9219 2023 Impact Factor: 5.6 2023 SCImago Journal Rankings: 2.340
ISI Accession Number ID	WOS:000622094200004

DC Field	Value	Language
dc.contributor.author	PENG, Y	-
dc.contributor.author	BAO, Y	-
dc.contributor.author	CHEN, Y	-
dc.contributor.author	Wu, C	-
dc.contributor.author	Meng, C	-
dc.contributor.author	Lin, W	-
dc.date.accessioned	2021-07-27T08:11:19Z	-
dc.date.available	2021-07-27T08:11:19Z	-
dc.date.issued	2021	-
dc.identifier.citation	IEEE Transactions on Parallel and Distributed Systems, 2021, v. 32 n. 8, p. 1947-1960	-
dc.identifier.issn	1045-9219	-
dc.identifier.uri	http://hdl.handle.net/10722/301454	-
dc.description.abstract	Efficient resource scheduling is essential for maximal utilization of expensive deep learning (DL) clusters. Existing cluster schedulers either are agnostic to machine learning (ML) workload characteristics, or use scheduling heuristics based on operators' understanding of particular ML framework and workload, which are less efficient or not general enough. In this article, we show that DL techniques can be adopted to design a generic and efficient scheduler. Specifically, we propose DL2, a DL-driven scheduler for DL clusters, targeting global training job expedition by dynamically resizing resources allocated to jobs. DL2 advocates a joint supervised learning and reinforcement learning approach: a neural network is warmed up via offline supervised learning based on job traces produced by the existing cluster scheduler; then the neural network is plugged into the live DL cluster, fine-tuned by reinforcement learning carried out throughout the training progress of the DL jobs, and used for deciding job resource allocation in an online fashion. We implement DL2 on Kubernetes and enable dynamic resource scaling in DL jobs on MXNet. Extensive evaluation shows that DL2 outperforms fairness scheduler (i.e., DRF) by 44.1 percent and expert heuristic scheduler (i.e., Optimus) by 17.5 percent in terms of average job completion time.	-
dc.language	eng	-
dc.publisher	Institute of Electrical and Electronics Engineers. The Journal's web site is located at http://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=71	-
dc.relation.ispartof	IEEE Transactions on Parallel and Distributed Systems	-
dc.rights	IEEE Transactions on Parallel and Distributed Systems. Copyright © Institute of Electrical and Electronics Engineers.	-
dc.rights	©20xx IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.	-
dc.subject	Deep learning	-
dc.subject	resource allocation	-
dc.subject	distributed training	-
dc.title	DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters	-
dc.type	Article	-
dc.identifier.email	Wu, C: cwu@cs.hku.hk	-
dc.identifier.authority	Wu, C=rp01397	-
dc.description.nature	link_to_OA_fulltext	-
dc.identifier.doi	10.1109/TPDS.2021.3052895	-
dc.identifier.scopus	eid_2-s2.0-85099732170	-
dc.identifier.hkuros	323504	-
dc.identifier.volume	32	-
dc.identifier.issue	8	-
dc.identifier.spage	1947	-
dc.identifier.epage	1960	-
dc.identifier.isi	WOS:000622094200004	-
dc.publisher.place	United States	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Article: DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats