{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:57:34Z","timestamp":1782845854121,"version":"3.54.5"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Software deployment is a critical software engineering practice, particularly for high performance computing (HPC) software.The deployment determines the software execution performance because the deployment maps a number of software components to multiple CPUs in a server. Inappropriate mapping decreases software parallelism and increases resource contention due to software co-location on the same CPUs. However, calculating the mapping to maximize the software performance is challenging, primarily due to the lack of a joint performance model that accounts for both software parallelism and co-location. Consequently, existing industry practice has to rely on experienced engineers to manually tune the mapping during deployment, resulting in substantial human resource waste of man-months and suboptimal software performance.<\/jats:p>\n                  <jats:p>This paper proposes a holistic approach to mapping multiple CPUs among multiple software components  \nto achieve better applicability and performance. We develop a performance model for predicting performance impact of different CPU mapping configurations, along with a search algorithm to identify the best mapping scheme. Our performance model jointly considers software parallelism and co-location, breaks the performance estimation into regularized execution and interference coefficient to improve accuracy, and integrates expert knowledge to reduce the model complexity. Our search algorithm employs nested iterative packing algorithm to explore all possible mapping schemes, thereby uncovering the optimal solution. Evaluation on a multi module HPC application shows 17% better performance than its default CPU mapping Our solution has been deployed in a commercial HPC cluster with more than 50K CPU cores, delivering 26.5% performance improvement and saving many man-months effort spent on performance tuning.<\/jats:p>","DOI":"10.1145\/3808183","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"4001-4024","source":"Crossref","is-referenced-by-count":0,"title":["Unleashing HPC Application Performance through Software Deployment: A Joint Model of Software Parallelism and Co-location"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-2678-9225","authenticated-orcid":false,"given":"Yuxin","family":"Ren","sequence":"first","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-7707-3468","authenticated-orcid":false,"given":"Li","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-0139-9055","authenticated-orcid":false,"given":"Chumin","family":"Sun","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-7059-5234","authenticated-orcid":false,"given":"Rui","family":"Fan","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-2553-1804","authenticated-orcid":false,"given":"Jie","family":"Sun","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0003-8246-4713","authenticated-orcid":false,"given":"Ning","family":"Jia","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-6909-9511","authenticated-orcid":false,"given":"Xinwei","family":"Hu","sequence":"additional","affiliation":[{"name":"Huawei Technologies, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","unstructured":"S. Balsamo A. Di Marco P. Inverardi and M. Simeoni. 2004. Model-based performance prediction in software development: a survey. IEEE Transactions on Software Engineering (2004). doi:10.1109\/TSE.2004.9 10.1109\/TSE.2004.9","DOI":"10.1109\/TSE.2004.9"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741962"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ECRTS.2009.14"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1137\/080738970"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3392717.3392764"},{"key":"e_1_2_1_6_1","unstructured":"CESM. [n. d.]. Community Earth System Model. https:\/\/2.zoppoz.workers.dev:443\/https\/www.cesm.ucar.edu\/"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00036"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00068"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243199"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2938172"},{"key":"e_1_2_1_11_1","unstructured":"Patent citation network. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/snap.stanford.edu\/data\/cit-Patents.html."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1137\/0209062"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540737"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1029\/2019ms001916"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451157"},{"key":"e_1_2_1_17_1","unstructured":"CESM Input Data. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/ftp.cgd.ucar.edu\/cesm\/inputdata\/."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522714"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694359"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 6th Conference on Symposium on Operating Systems Design and Implementation (OSDI'04)","author":"Dean Jeffrey","year":"2004","unstructured":"Jeffrey Dean and Sanjay Ghemawat. 2004. MapReduce: Simplified Data Processing on Large Clusters. In Proceedings of the 6th Conference on Symposium on Operating Systems Design and Implementation (OSDI'04)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS62706.2024.00034"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3337821.3337893"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2018.00050"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416620"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.11"},{"key":"e_1_2_1_26_1","unstructured":"Twitter follower network. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/snap.stanford.edu\/data\/twitter-2010.html."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2013.6606597"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18).","author":"Funston Justin","year":"2018","unstructured":"Justin Funston, Maxime Lorrillere, Alexandra Fedorova, Baptiste Lepers, David Vengerov, Jean-Pierre Lozi, and Vivien Qu\u00e9ma. 2018. Placement of Virtual Containers on NUMA Systems: A Practical and Comprehensive Model. In Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2019.00015"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3064176"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2013.6693089"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","unstructured":"Qianyu Guo Sen Chen Xiaofei Xie Lei Ma Qiang Hu Hongtao Liu Yang Liu Jianjun Zhao and Xiaohong Li. 2019. An Empirical Study Towards Characterizing Deep Learning Development and Deployment Across Different Frameworks and Platforms. In 2019 34th IEEE\/ACM International Conference on Automated Software Engineering (ASE'19). doi:10.1109\/ASE.2019.00080 10.1109\/ASE.2019.00080","DOI":"10.1109\/ASE.2019.00080"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2592798.2592807"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764800"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1175\/BAMS-D-12-00121.1"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18).","author":"Iorgulescu C\u0103lin","year":"2018","unstructured":"C\u0103lin Iorgulescu, Reza Azimi, Youngjin Kwon, Sameh Elnikety, Manoj Syamala, Vivek Narasayya, Herodotos Herodotou, Paulo Tomita, Alex Chen, Jack Zhang, and Junhua Wang. 2018. PerfIso: Performance Isolation for Commercial Latency- Sensitive Services. In Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference (USENIX ATC'18)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion58688"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-Companion.2019.00028"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400704"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807632"},{"key":"e_1_2_1_42_1","volume-title":"Thread and Memory Placement on NUMA Systems: Asymmetry Matters. In 2015 USENIX Annual Technical Conference (USENIX ATC'15)","author":"Lepers Baptiste","year":"2015","unstructured":"Baptiste Lepers, Vivien Quema, and Alexandra Fedorova. 2015. Thread and Memory Placement on NUMA Systems: Asymmetry Matters. In 2015 USENIX Annual Technical Conference (USENIX ATC'15)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307681.3325409"},{"key":"e_1_2_1_44_1","unstructured":"Lightweight Python library for in-memory matrix completion. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/tonyduan\/matrix-completion."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.75"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.84"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815406"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1294904.1294911"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/103727.103729"},{"key":"e_1_2_1_50_1","unstructured":"The Community Earth System Model. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/ESCOMP\/CESM."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2019.00065"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00176"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-4380-9_2"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/1048935.1050204"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824043"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/3015274.3015275"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254082"},{"key":"e_1_2_1_58_1","volume-title":"Multi-tenant Edge Clouds. In 2020 USENIX Annual Technical Conference (USENIX ATC'20)","author":"Ren Yuxin","year":"2020","unstructured":"Yuxin Ren, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood, and Alain Tchana. 2020. Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds. In 2020 USENIX Annual Technical Conference (USENIX ATC'20)."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2018.00025"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3361525.3361537"},{"key":"e_1_2_1_61_1","unstructured":"California road network. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/snap.stanford.edu\/data\/roadNet-CA.html."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132771"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3392717.3392765"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370833"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/2889160.2889223"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2786805.2786845"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2012.6227196"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3431379.3460635"},{"key":"e_1_2_1_69_1","unstructured":"Phoronix Test Suite. [n. d.]. https:\/\/2.zoppoz.workers.dev:443\/https\/www.phoronix-test-suite.com\/."},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508244.1508274"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356152"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS59052.2023.00030"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/1088149.1088190"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00100"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00072"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446083"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504189"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00099"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882362.1882445"},{"key":"e_1_2_1_80_1","doi-asserted-by":"crossref","first-page":"977","DOI":"10.5194\/gmd-13-977-2020","article-title":"Beijing Climate Center Earth System Model version 1 (BCC-ESM1): model description and evaluation of aerosol simulations","volume":"13","author":"Wu T.","year":"2020","unstructured":"T. Wu, F. Zhang, J. Zhang, W. Jie, Y. Zhang, F. Wu, L. Li, J. Yan, X. Liu, X. Lu, H. Tan, L. Zhang, J. Wang, and A. Hu. 2020. Beijing Climate Center Earth System Model version 1 (BCC-ESM1): model description and evaluation of aerosol simulations. Geoscientific Model Development 13, 3 (2020), 977-1005. https:\/\/2.zoppoz.workers.dev:443\/https\/gmd.copernicus.org\/articles\/13\/977\/2020\/","journal-title":"Geoscientific Model Development"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP.2019.00010"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"},{"key":"e_1_2_1_83_1","unstructured":"Lei Zhang Takayuki Okamoto Shuuichirou Ishii Kouichi Hirai Shinji Sumimoto Balazs Gerofi Masamichi Takagi and Yutaka Ishikawa. [n. d.]. OS Enhancement in Supercomputer Fugaku https:\/\/2.zoppoz.workers.dev:443\/https\/www.fujitsu.com\/global\/about\/ resources\/publications\/technicalreview\/2020-03\/article06.html."},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2015.15"},{"key":"e_1_2_1_85_1","volume-title":"Gemini: A Computation-Centric Distributed Graph Processing System. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI'16)","author":"Zhu Xiaowei","year":"2016","unstructured":"Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma. 2016. Gemini: A Computation-Centric Distributed Graph Processing System. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI'16)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3808183","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:57:03Z","timestamp":1782842223000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3808183"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":85,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808183"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3808183","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}