{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T12:02:58Z","timestamp":1785240178827,"version":"3.55.0"},"reference-count":30,"publisher":"Wiley","issue":"16","license":[{"start":{"date-parts":[[2019,10,21]],"date-time":"2019-10-21T00:00:00Z","timestamp":1571616000000},"content-version":"am","delay-in-days":365,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#am"},{"start":{"date-parts":[[2018,10,21]],"date-time":"2018-10-21T00:00:00Z","timestamp":1540080000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"publisher","award":["DE-AC02-05CH11231"],"award-info":[{"award-number":["DE-AC02-05CH11231"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2019,8,25]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>Deep learning has proven to be a successful tool for solving a large variety of problems in various scientific fields and beyond. In recent years, the models as well as the available datasets have grown bigger and more complicated, and thus, an increasing amount of computing resources is required in order to train these models in a reasonable amount of time. Besides being able to use HPC resources, deep learning model developers want flexible frameworks which allow for rapid prototyping. One of the most important of these frameworks is Google TensorFlow, which provides both features, ie, good performance as well as flexibility. In this paper, we discuss different solutions for scaling the TensorFlow Framework to thousands of nodes on contemporary Cray XC supercomputing systems.<\/jats:p>","DOI":"10.1002\/cpe.4989","type":"journal-article","created":{"date-parts":[[2018,10,28]],"date-time":"2018-10-28T20:54:54Z","timestamp":1540760094000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["TensorFlow at Scale: Performance and productivity analysis of distributed training with Horovod, MLSL, and Cray PE ML"],"prefix":"10.1002","volume":"31","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-0832-6198","authenticated-orcid":false,"given":"Thorsten","family":"Kurth","sequence":"first","affiliation":[{"name":"National Energy Research Scientific Computing Center Lawrence Berkeley National Laboratory  Berkeley California"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mikhail","family":"Smorkalov","sequence":"additional","affiliation":[{"name":"Software and Services Group Intel Corporation  Moscow Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter","family":"Mendygral","sequence":"additional","affiliation":[{"name":"Cray Programming Environments Performance Engineering Cray Inc  Bloomington Minnesota"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Srinivas","family":"Sridharan","sequence":"additional","affiliation":[{"name":"Parallel Computing Labs Intel Corporation  Karnataka India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amrita","family":"Mathuriya","sequence":"additional","affiliation":[{"name":"Data Center Group Intel Corporation  Hillsboro Oregon"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2018,10,21]]},"reference":[{"key":"e_1_2_10_2_1","unstructured":"The message passing interface (MPI) standard.https:\/\/2.zoppoz.workers.dev:443\/http\/www.mcs.anl.gov\/research\/projects\/mpi\/index.htm"},{"key":"e_1_2_10_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/0925-2312(93)90006-O"},{"key":"e_1_2_10_4_1","unstructured":"AbadiM AgarwalA BarhamP et al.TensorFlow: large\u2010scale machine learning on heterogeneous systems software.2015."},{"key":"e_1_2_10_5_1","unstructured":"TensorFlow.https:\/\/2.zoppoz.workers.dev:443\/https\/tensorflow.org"},{"key":"e_1_2_10_6_1","unstructured":"PirogovV GennadyF.Introducing DNN primitives in Intel\u00aeMath Kernel Library.2016.https:\/\/2.zoppoz.workers.dev:443\/https\/software.intel.com\/en-us\/articles\/introducing-dnn-primitives-in-intelr-mkl"},{"key":"e_1_2_10_7_1","unstructured":"ChetlurS WoolleyC VandermerschP et al.cuDNN: efficient primitives for deep learning.2014; arXiv preprint arXiv:1410.0759."},{"key":"e_1_2_10_8_1","unstructured":"NVIDIA cuDNN GPU accelerated deep learning.https:\/\/2.zoppoz.workers.dev:443\/https\/developer.nvidia.com\/cudnn"},{"key":"e_1_2_10_9_1","unstructured":"GRPC.https:\/\/2.zoppoz.workers.dev:443\/https\/grpc.io"},{"key":"e_1_2_10_10_1","unstructured":"NERSC Cori.https:\/\/2.zoppoz.workers.dev:443\/https\/www.nersc.gov\/users\/computational-systems\/cori\/"},{"key":"e_1_2_10_11_1","unstructured":"MathuriyaA KurthT RaneV et al.Scaling GRPC TensorFlow on 512 nodes of cori supercomputer.2017; arXiv e\u2010prints."},{"key":"e_1_2_10_12_1","unstructured":"SergeevA Del BalsoM.Horovod: fast and easy distributed deep learning in TensorFlow.2018; arXiv preprint arXiv:1802.05799."},{"key":"e_1_2_10_13_1","unstructured":"SergeevA.Horovod \u2010 distributed TensorFlow made easy.2017.https:\/\/2.zoppoz.workers.dev:443\/https\/www.slideshare.net\/AlexanderSergeev4\/horovod-distributed-tensorflow-made-easy"},{"key":"e_1_2_10_14_1","unstructured":"SergeevA Del BalsoM.Meet Horovod: Uber's open source distributed deep learning framework for TensorFlow.2017;https:\/\/2.zoppoz.workers.dev:443\/https\/eng.uber.com\/horovod\/"},{"key":"e_1_2_10_15_1","unstructured":"Open MPI: open source high performance computing.https:\/\/2.zoppoz.workers.dev:443\/https\/www.open-mpi.org"},{"key":"e_1_2_10_16_1","unstructured":"SridharanS VaidyanathanK KalamkarD et al.On scale\u2010out deep learning training for cloud and HPC. Paper presented at: SysML 2018 Conference;2018;Stanford CA."},{"key":"e_1_2_10_17_1","unstructured":"Intel\u00aeMachine Learning Scaling Library for Linux* OS.2017.https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/01org\/MLSL"},{"key":"e_1_2_10_18_1","unstructured":"DeepBench.https:\/\/2.zoppoz.workers.dev:443\/https\/svail.github.io\/DeepBench\/"},{"key":"e_1_2_10_19_1","unstructured":"Adding a new OP.https:\/\/2.zoppoz.workers.dev:443\/https\/www.tensorflow.org\/extend\/adding_an_op"},{"key":"e_1_2_10_20_1","doi-asserted-by":"crossref","unstructured":"BhimjiW FarrellSA KurthT PaganiniM Prabhat RacahE.Deep neural networks for physics analysis on low\u2010level whole\u2010detector data at the LHC. Paper presented at: ACAT 2017 Conference;2017;Seattle WA.","DOI":"10.1088\/1742-6596\/1085\/4\/042034"},{"key":"e_1_2_10_21_1","doi-asserted-by":"crossref","unstructured":"KurthT ZhangJ SatishN et al.Deep learning at 15PF: supervised and semi\u2010supervised classification for scientific data.2017; arXiv:1708.05256 [cs.PF].","DOI":"10.1145\/3126908.3126916"},{"key":"e_1_2_10_22_1","unstructured":"HEP\u2010CNN benchmark repository.https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/NERSC\/hep_cnn_benchmark"},{"key":"e_1_2_10_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/JHEP02(2014)057"},{"key":"e_1_2_10_24_1","unstructured":"MustafaM BardD BhimjiW Al\u2010RfouR Luki\u0107Z.Creating virtual universes using generative adversarial networks.2017; arXiv e\u2010prints."},{"key":"e_1_2_10_25_1","unstructured":"Bazel.https:\/\/2.zoppoz.workers.dev:443\/https\/bazel.build"},{"key":"e_1_2_10_26_1","unstructured":"Piz Daint.https:\/\/2.zoppoz.workers.dev:443\/https\/www.cscs.ch\/computers\/piz-daint\/"},{"key":"e_1_2_10_27_1","unstructured":"NCCL.https:\/\/2.zoppoz.workers.dev:443\/https\/developer.nvidia.com\/nccl"},{"key":"e_1_2_10_28_1","unstructured":"KingmaDP BaJ.Adam: a method for stochastic optimization.2014; arXiv:1412.6980."},{"key":"e_1_2_10_29_1","unstructured":"ChilimbiT SuzueY ApacibleJ KalyanaramanK.Project Adam: building an efficient and scalable deep learning training system. In: Proceedings of the 11th USENIX Symposium on Operating Systems Design and Implementation;2014;Broomfield CO."},{"key":"e_1_2_10_30_1","unstructured":"YouY GitmanI GinsburgB.Large batch training of convolutional networks.2017; arXiv e\u2010prints."},{"key":"e_1_2_10_31_1","unstructured":"Summit.https:\/\/2.zoppoz.workers.dev:443\/https\/www.olcf.ornl.gov\/olcf-resources\/compute-systems\/summit\/"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fcpe.4989","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.4989","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.4989","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/am-pdf\/10.1002\/cpe.4989","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.4989","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T19:01:08Z","timestamp":1694286068000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.4989"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,21]]},"references-count":30,"journal-issue":{"issue":"16","published-print":{"date-parts":[[2019,8,25]]}},"alternative-id":["10.1002\/cpe.4989"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/cpe.4989","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,10,21]]},"assertion":[{"value":"2018-06-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-08-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-10-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e4989"}}