{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T23:17:28Z","timestamp":1771629448254,"version":"3.50.1"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Union\u2019s Horizon 2020 research and innovation programme","award":["899546"],"award-info":[{"award-number":["899546"]}]},{"name":"The Coordena\u00e7\u00e3o de Aperfei\u00e7oamento de Pessoal de N\u00edvel Superior - Brazil"},{"name":"Indam GNCS Project 2022","award":["CUP_E55F22000270001"],"award-info":[{"award-number":["CUP_E55F22000270001"]}]},{"name":"European Union under NextGenerationEU"},{"name":"ChipIR","award":["ISIS.E.RB2200004-1, 10.5286\/ISIS.E.RB2000137-1, 10.5286\/ISIS.E.101136531"],"award-info":[{"award-number":["ISIS.E.RB2200004-1, 10.5286\/ISIS.E.RB2000137-1, 10.5286\/ISIS.E.101136531"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>Graphics Processing Units (GPUs) compilers have evolved in order to support general-purpose programming languages for multiple architectures. NVIDIA CUDA Compiler (NVCC) has many compilation levels before generating the machine code and applies complex optimizations to improve performance. These optimizations modify how the software is mapped in the underlying hardware; thus, as we show in this article, they can also affect GPU reliability. We evaluate the effects on the GPU error rate of the optimization flags applied at the NVCC Parallel Thread Execution (PTX) compiling phase by analyzing two NVIDIA GPU architectures (Kepler and Volta) and two compiler versions (NVCC 10.2 and 11.3). We compare and combine fault propagation analysis based on software fault injection, hardware utilization distribution obtained with application-level profiling, and machine instructions radiation-induced error rate measured with beam experiments. We consider eight different workloads and 144 combinations of compilation flags, and we show that optimizations can impact the GPUs\u2019 error rate of up to an order of magnitude. Additionally, through accelerated neutron beam experiments on a NVIDIA Kepler GPU, we show that the error rate of the unoptimized GEMM (-O0 flag) is lower than the optimized GEMM\u2019s (-O3 flag) error rate. When the performance is evaluated together with the error rate, we show that the most optimized versions (-O1 and -O3) always produce a higher amount of correct data than the unoptimized code (-O0).<\/jats:p>","DOI":"10.1145\/3638249","type":"journal-article","created":{"date-parts":[[2024,1,12]],"date-time":"2024-01-12T11:15:46Z","timestamp":1705058146000},"page":"1-22","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Assessing the Impact of Compiler Optimizations on GPUs Reliability"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-3504-9862","authenticated-orcid":false,"given":"Fernando Fernandes Dos","family":"Santos","sequence":"first","affiliation":[{"name":"Univ Rennes, INRIA, Rennes, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-7402-4780","authenticated-orcid":false,"given":"Luigi","family":"Carro","sequence":"additional","affiliation":[{"name":"Institute of Informatics, Federal University of Rio Grande do Sul, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-5676-9228","authenticated-orcid":false,"given":"Flavio","family":"Vella","sequence":"additional","affiliation":[{"name":"University of Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-0821-1879","authenticated-orcid":false,"given":"Paolo","family":"Rech","sequence":"additional","affiliation":[{"name":"University of Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,2,15]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"15","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Anwer Abdul Rehman","year":"2020","unstructured":"Abdul Rehman Anwer, Guanpeng Li, Karthik Pattabiraman, Michael Sullivan, Timothy Tsai, and Siva Kumar Sastry Hari. 2020. GPU-trident: Efficient modeling of error propagation in GPU programs. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Atlanta, Georgia). IEEE Press, 15 pages."},{"key":"e_1_3_2_3_2","first-page":"1274","volume-title":"Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium Workshops","author":"Ashraf R. A.","year":"2017","unstructured":"R. A. Ashraf, R. Gioiosa, G. Kestor, and R. F. DeMara. 2017. Exploring the effect of compiler optimizations on the reliability of HPC applications. In Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium Workshops. 1274\u20131283."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","unstructured":"J. M. Badia G. Leon J. A. Belloch M. Garcia-Valderas A. Lindoso and L. Entrena. 2022. Comparison of parallel implementation strategies in GPU-accelerated system-on-chip under proton irradiation. In IEEE Transactions on Nuclear Science 69 3 (2022) 444\u2013452. DOI:10.1109\/TNS.2021.3128722","DOI":"10.1109\/TNS.2021.3128722"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2005.69"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMSCS.2018.2797195"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3128501"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Carlo Cazzaniga and Christopher D. Frost. 2018. Progress of the scientific commissioning of a fast neutron beamline for chip irradiation. In Journal of Physics: Conference Series 1021 1 (2018) 012037.","DOI":"10.1088\/1742-6596\/1021\/1\/012037"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 2019 49th Annual IEEE\/IFIP International Conference on Dependable Systems and Networks","author":"Chatzidimitriou A.","year":"2019","unstructured":"A. Chatzidimitriou, P. Bodmann, G. Papadimitriou, D. Gizopoulos, and P. Rech. 2019. Demystifying soft error assessment strategies on arm CPUs: Microarchitectural fault injection vs. neutron beam experiments. In Proceedings of the 2019 49th Annual IEEE\/IFIP International Conference on Dependable Systems and Networks. 26\u201338."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/dsn-w.2017.16"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356177"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488859"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","unstructured":"Josie E. Rodriguez Condia Boyang Du Matteo Sonza Reorda and Luca Sterpone. 2020. FlexGripPlus: An improved GPGPU model to support reliability analysis. Microelectronics Reliability 109 (2020) 113660. 10.1016\/j.microrel.2020.113660","DOI":"10.1016\/j.microrel.2020.113660"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/2354410.2355150"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2014.6844486"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/1735971.1736063"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","unstructured":"F. Fernandes dos Santos C. Lunardi D. Oliveira F. Libano and P. Rech. 2019. Reliability evaluation of mixed-precision architectures. In 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA\u201919). Washington DC USA 238\u2013249. DOI:10.1109\/HPCA.2019.00041","DOI":"10.1109\/HPCA.2019.00041"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10836-015-5555-z"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2021.3119511"},{"issue":"3","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"791","DOI":"10.1109\/TC.2015.2444855","article-title":"Evaluation and mitigation of radiation-induced soft errors in graphics processing units","volume":"65","author":"Oliveira D. A. G. Goncalves de","year":"2016","unstructured":"D. A. G. Goncalves de Oliveira, L. L. Pilla, T. Santini, and P. Rech. 2016. Evaluation and mitigation of radiation-induced soft errors in graphics processing units. IEEE Transactions on Computers 65, 3 (2016), 791\u2013804.","journal-title":"IEEE Transactions on Computers"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2012.6263960"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TDSC.2021.3063083"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2014.6853212"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2021.3098845"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/dft.2016.7684076"},{"key":"e_1_3_2_27_2","volume-title":"Measurement and Reporting of Alpha Particle and Terrestrial Cosmic Ray-Induced Soft Errors in Semiconductor Devices","year":"2006","unstructured":"JEDEC. 2006. Measurement and Reporting of Alpha Particle and Terrestrial Cosmic Ray-Induced Soft Errors in Semiconductor Devices. Technical Report JESD89A. JEDEC Standard."},{"key":"e_1_3_2_28_2","unstructured":"Saurabh Jha Timothy Tsai Siva Hari Michael Sullivan Zbigniew Kalbarczyk Stephen W. Keckler and Ravishankar K. Iyer. 2019. Kayotee: A Fault Injection-based System to Assess the Safety and Reliability of Autonomous Vehicles to Faults and Errors. arXiv:1907.01024. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/1907.01024"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3434402"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.microrel.2020.113856"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2018.00038"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2018.2884460"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2018.2823786"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2011.2171993"},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the 2018 International Conference on Field-Programmable Technology","author":"Marty Thibaut","year":"2018","unstructured":"Thibaut Marty, Tomofumi Yuki, and Steven Derrien. 2018. Enabling overclocking through algorithm-level error detection. In Proceedings of the 2018 International Conference on Field-Programmable Technology."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956570"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126960"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/DFT.2014.6962085"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2008.917757"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2019.00024"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC53511.2021.00021"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00075"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1007\/978-3-642-38853-8_5","volume-title":"Proceedings of the Embedded Systems: Design, Analysis and Verification.","author":"Parizi Rafael B.","year":"2013","unstructured":"Rafael B. Parizi, Ronaldo R. Ferreira, Luigi Carro, and \u00c1lvaro F. Moreira. 2013. Compiler optimizations do impact the reliability of control-flow radiation hardened embedded software. In Proceedings of the Embedded Systems: Design, Analysis and Verification.Gunar Schirner, Marcelo G\u00f6tz, Achim Rettberg, Mauro C. Zanella, and Franz J. Rammig (Eds.), Springer, Berlin, 49\u201360."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISIE.2008.4677290"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/MWSCAS.2017.8053069"},{"key":"e_1_3_2_46_2","volume-title":"Proceedings of the IEEE 10th Workshop on Silicon Errors in Logic - System Effects","author":"Rech Paolo","year":"2014","unstructured":"Paolo Rech, Luigi Carro, Nicholas Wang, Timothy Tsai, Siva Kumar Sastry Hari, and Stephen W. Keckler. 2014. Measuring the radiation reliability of SRAM structures in GPUS designed for HPC. In Proceedings of the IEEE 10th Workshop on Silicon Errors in Logic - System Effects."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2013.2286970"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.21"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00037"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2018.2878387"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2009.4798243"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480111"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN48987.2021.00041"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2016.7482077"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","unstructured":"Alessandro Vallero Dimitris Gizopoulos and Stefano Di Carlo. 2017. SIFI: AMD southern islands GPU microarchitectural level fault injector. In Proceedings of the 2017 IEEE 23rd International Symposium on On-Line Testing and Robust System Design. 138\u2013144. DOI:DOI:10.1109\/IOLTS.2017.8046209","DOI":"10.1109\/IOLTS.2017.8046209"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.5555\/2665671.2665686"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00114"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.36"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3638249","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3638249","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:53:35Z","timestamp":1750287215000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3638249"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,15]]},"references-count":57,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3638249"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3638249","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,15]]},"assertion":[{"value":"2023-06-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-04","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}