{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:58:03Z","timestamp":1782845883811,"version":"3.54.5"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Code language models have demonstrated strong capabilities across a wide range of code intelligence tasks. While the majority of existing research prioritizes performance improvements on benchmark datasets, few of them have focused on the internal interpretability of models\u2014how specific neurons affect linguistic features such as syntax and semantics, which is critical for model transparency, controllability, and reliability. Although various neuron interpretability techniques have been developed in NLP, directly applying them to source code yields suboptimal results due to the unique characteristics of programming languages, such as their formal structure, hierarchical organization, and executability.  \nIn this work, we empirically investigate the intrinsic mechanisms of code LLMs at the neuron level, aiming to localize both language-specific neurons (i.e., neurons that are selectively responsive to individual programming languages) and concept layers (i.e., feed-forward layers that encode language-agnostic representations of code).  \nOur study employs two state-of-the-art models, Llama-3.1-8B and Qwen2.5-Coder-32B, across five programming languages: C++, Java, Python, Go, and JavaScript. By analyzing neuron activation patterns in response to multilingual code inputs, we investigate the role of individual neurons and the contribution of different layers during output generation.  \nOur empirical findings reveal that: (1) code LLMs contain neurons specialized for individual programming languages, alongside a universal subset that supports general-purpose code generation; and (2) lower layers primarily encode language-specific syntactic structures, while middle layers capture semantic abstractions that generalize across languages, manifesting as concept layers.  \nTo demonstrate the practical usability of these findings, we apply our findings to three downstream tasks: neuron-guided fine-tuning for code generation, clone detection using concept-layer embeddings, and transfer learning guided by concept-layer representations for code summarization. Experimental evaluations show that each strategy consistently improves the performance of multilingual code LLMs.<\/jats:p>","DOI":"10.1145\/3797083","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"1219-1240","source":"Crossref","is-referenced-by-count":0,"title":["Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0005-1589-6254","authenticated-orcid":false,"given":"Zhe","family":"Yin","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-0529-6408","authenticated-orcid":false,"given":"Xiaodong","family":"Gu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-8370-3956","authenticated-orcid":false,"given":"Beijun","family":"Shen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025","author":"Chai Linzheng","year":"2025","unstructured":"Linzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin, Ke Jin, Jiaheng Liu, Tao Sun, Ge Zhang, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang, Xianjie Wu, Bing Wang, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang, and Zhoujun Li. 2025. McEval: Massively Multilingual Code Evaluation. In Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025."},{"key":"e_1_2_1_2_1","volume-title":"Software Engineering","author":"Cito J\u00fcrgen","year":"2024","unstructured":"J\u00fcrgen Cito, Isil Dillig, Vijayaraghavan Murali, and Satish Chandra. 2024. Counterfactual Explanations for Models of Code. In Software Engineering 2024, Fachtagung des GI-Fachbereichs Softwaretechnik, Linz, Austria, February 26 -March 1, 2024 (LNI, Vol. P-343). Gesellschaft f\u00fcr Informatik e.V., 91-92."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL 2019","author":"Clark Kevin","year":"2019","unstructured":"Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What Does BERT Look at? An Analysis of BERT's Attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@ACL 2019, Florence, Italy, August 1, 2019. Association for Computational Linguistics, 276-286."},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019","author":"Conneau Alexis","year":"2019","unstructured":"Alexis Conneau and Guillaume Lample. 2019. Cross-lingual Language Model Pretraining. In Proceedings of Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. 7057-7067."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.581"},{"key":"e_1_2_1_6_1","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman and et al. 2024. The Llama 3 Herd of Models. CoRR abs\/2407.21783 (2024)."},{"key":"e_1_2_1_7_1","volume-title":"Toy Models of Superposition. CoRR abs\/2209.10652","author":"Elhage Nelson","year":"2022","unstructured":"Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger B. Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy Models of Superposition. CoRR abs\/2209.10652 (2022)."},{"key":"e_1_2_1_8_1","volume-title":"CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL","volume":"1547","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020). Association for Computational Linguistics, 1536-1547."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of Forty-first International Conference on Machine Learning, ICML 2024","author":"Ghandeharioun Asma","year":"2024","unstructured":"Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024. Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models. In Proceedings of Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024."},{"key":"e_1_2_1_10_1","first-page":"6779","volume-title":"Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 -Volume 1: Long Papers","author":"Gupta Mukur","year":"2025","unstructured":"Mukur Gupta, Noopur Bhatt, and Suman Jana. 2025. CodeSCM: Causal Analysis for Multi-Modal Code Generation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 -Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 -May 4, 2025. Association for Computational Linguistics, 6779-6793."},{"key":"e_1_2_1_11_1","volume-title":"Finding Neurons in a Haystack: Case Studies with Sparse Probing. Trans. Mach. Learn. Res. 2023","author":"Gurnee Wes","year":"2023","unstructured":"Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023. Finding Neurons in a Haystack: Case Studies with Sparse Probing. Trans. Mach. Learn. Res. 2023 (2023)."},{"key":"e_1_2_1_12_1","volume-title":"Looking into Black Box Code Language Models. CoRR abs\/2407.04868","author":"Haider Muhammad Umair","year":"2024","unstructured":"Muhammad Umair Haider, Umar Farooq, A. B. Siddique, and Mark Marron. 2024. Looking into Black Box Code Language Models. CoRR abs\/2407.04868 (2024)."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019","author":"Hewitt John","year":"2019","unstructured":"John Hewitt and Percy Liang. 2019. Designing and Interpreting Probes with Control Tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019. Association for Computational Linguistics, 2733-2743."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019","volume":"1","author":"Hewitt John","year":"2019","unstructured":"John Hewitt and Christopher D. Manning. 2019. A Structural Probe for Finding Syntax in Word Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers). Association for Computational Linguistics, 4129-4138."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of Forty-first International Conference on Machine Learning, ICML 2024","author":"Hooda Ashish","year":"2024","unstructured":"Ashish Hooda, Mihai Christodorescu, Miltiadis Allamanis, Aaron Wilson, Kassem Fawaz, and Somesh Jha. 2024. Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates. In Proceedings of Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024."},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of 2026 ACM International Conference on the Foundations of Software Engineering (FSE). ACM.","author":"Hu Chao","year":"2026","unstructured":"Chao Hu, Wenhao Zeng, Yuling Shi, Beijun Shen, and Xiaodong Gu. 2026. In Line with Context: Repository-Level Code Generation via Context Inlining. In Proceedings of 2026 ACM International Conference on the Foundations of Software Engineering (FSE). ACM."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the Tenth International Conference on Learning Representations, ICLR 2022","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022."},{"key":"e_1_2_1_18_1","volume-title":"CoRR abs\/2409.12186","author":"Hui Binyuan","year":"2024","unstructured":"Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Kai Dang, An Yang, Rui Men, Fei Huang, Xingzhang Ren, Xuancheng Ren, Jingren Zhou, and Junyang Lin. 2024. Qwen2.5-Coder Technical Report. CoRR abs\/2409.12186 (2024)."},{"key":"e_1_2_1_19_1","volume-title":"CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs\/1909.09436","author":"Husain Hamel","year":"2019","unstructured":"Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. CoRR abs\/1909.09436 (2019)."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1356"},{"key":"e_1_2_1_21_1","first-page":"432","volume-title":"Proceedings of 18th IEEE\/ACM International Conference on Mining Software Repositories, MSR 2021","author":"Jiarpakdee Jirayus","year":"2021","unstructured":"Jirayus Jiarpakdee, Chakkrit Tantithamthavorn, and John C. Grundy. 2021. Practitioners' Perceptions of the Goals and Visual Explanations of Defect Prediction Models. In Proceedings of 18th IEEE\/ACM International Conference on Mining Software Repositories, MSR 2021, Madrid, Spain, May 17-19, 2021. 432-443."},{"key":"e_1_2_1_22_1","volume-title":"ACL 2025","author":"Kargaran Amir Hossein","year":"2025","unstructured":"Amir Hossein Kargaran, Yihong Liu, Fran\u00e7ois Yvon, and Hinrich Sch\u00fctze. 2025. How Programming Concepts and Neurons Are Shared in Code Language Models. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 -August 1, 2025. Association for Computational Linguistics, 26905-26917."},{"key":"e_1_2_1_23_1","volume-title":"Overcoming catastrophic forgetting in neural networks. CoRR abs\/1612.00796","author":"Kirkpatrick James","year":"2016","unstructured":"James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016. Overcoming catastrophic forgetting in neural networks. CoRR abs\/1612.00796 (2016)."},{"key":"e_1_2_1_24_1","volume-title":"Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience 2","author":"Kriegeskorte Nikolaus","year":"2008","unstructured":"Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. 2008. Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience 2 (2008)."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, NeurIPS 2024","author":"Li Linyi","year":"2024","unstructured":"Linyi Li, Shijie Geng, Zhenwen Li, Yibo He, Hao Yu, Ziyue Hua, Guanghan Ning, Siwei Wang, Tao Xie, and Hongxia Yang. 2024. InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models. In Proceedings of Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems, NeurIPS 2024, Vancouver, BC, Canada, December 10 -15, 2024."},{"key":"e_1_2_1_26_1","article-title":"On the Reliability and Explainability of Language Models for Program Generation","volume":"33","author":"Liu Yue","year":"2024","unstructured":"Yue Liu, Chakkrit Tantithamthavorn, Yonghui Liu, and Li Li. 2024. On the Reliability and Explainability of Language Models for Program Generation. ACM Trans. Softw. Eng. Methodol. 33, 5 (2024), 126:1-126:26.","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of 7th International Conference on Learning Representations, ICLR 2019","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In Proceedings of 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019."},{"key":"e_1_2_1_28_1","first-page":"13794","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023","author":"Ma Xinyu","year":"2023","unstructured":"Xinyu Ma, Xuebo Liu, and Min Zhang. 2023. Clustering Pseudo Language Family in Multilingual Translation Models with Fisher Information Matrix. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023. Association for Computational Linguistics, 13794-13804."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the Twelfth International Conference on Learning Representations, ICLR 2024","author":"Merullo Jack","year":"2024","unstructured":"Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024. Circuit Component Reuse Across Tasks in Transformer Language Models. In Proceedings of the Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024."},{"key":"e_1_2_1_30_1","first-page":"96","volume-title":"Proceedings of 23rd IEEE International Working Conference on Source Code Analysis and Manipulation, SCAM 2023","author":"Mohammadkhani Ahmad Haji","year":"2023","unstructured":"Ahmad Haji Mohammadkhani, Chakkrit Tantithamthavorn, and Hadi Hemmati. 2023. Explaining Transformer-based Code Models: What Do They Learn? When They Do Not Work?. In Proceedings of 23rd IEEE International Working Conference on Source Code Analysis and Manipulation, SCAM 2023, Bogot\u00e1, Colombia, October 2-3, 2023. IEEE, 96-106."},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023","author":"Morris John X.","year":"2023","unstructured":"John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush. 2023. Text Embeddings Reveal (Almost) As Much As Text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023. Association for Computational Linguistics, 12448-12460."},{"key":"e_1_2_1_32_1","first-page":"1023","volume-title":"Proc. ACM Softw. Eng. 2, FSE","author":"Moumoula Micheline B\u00e9n\u00e9dicte","year":"2025","unstructured":"Micheline B\u00e9n\u00e9dicte Moumoula, Abdoul Kader Kabor\u00e9, Jacques Klein, and Tegawend\u00e9 F. Bissyand\u00e9. 2025. The Struggles of LLMs in Cross-Lingual Code Clone Detection. Proc. ACM Softw. Eng. 2, FSE (2025), 1023-1045."},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","first-page":"1215","DOI":"10.1109\/TSE.2024.3379943","article-title":"Toward a Theory of Causation for Interpreting Neural Code Models","volume":"50","author":"Nader-Palacio David","year":"2024","unstructured":"David Nader-Palacio, Alejandro Velasco, Nathan Cooper, Alvaro Rodriguez, Kevin Moran, and Denys Poshyvanyk. 2024. Toward a Theory of Causation for Interpreting Neural Code Models. IEEE Trans. Software Eng. 50, 5 (2024), 1215-1243.","journal-title":"IEEE Trans. Software Eng."},{"key":"e_1_2_1_34_1","first-page":"867","volume-title":"Proceedings of 36th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2021","author":"Paltenghi Matteo","year":"2021","unstructured":"Matteo Paltenghi and Michael Pradel. 2021. Thinking Like a Developer? Comparing the Attention of Humans with Neural Models of Code. In Proceedings of 36th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2021, Melbourne, Australia, November 15-19, 2021. IEEE, 867-879."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. 8024-8035."},{"key":"e_1_2_1_36_1","first-page":"3138","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020","author":"Pimentel Tiago","year":"2020","unstructured":"Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020. Pareto Probing: Trading Off Accuracy for Complexity. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020. Association for Computational Linguistics, 3138-3153."},{"key":"e_1_2_1_37_1","first-page":"512","volume-title":"Proceedings of IEEE International Conference on Software Maintenance and Evolution, ICSME 2024","author":"Pinku Subroto Nag","year":"2024","unstructured":"Subroto Nag Pinku, Debajyoti Mondal, and Chanchal K. Roy. 2024. On the Use of Deep Learning Models for Semantic Clone Detection. In Proceedings of IEEE International Conference on Software Maintenance and Evolution, ICSME 2024, Flagstaff, AZ, USA, October 6-11, 2024. IEEE, 512-524."},{"key":"e_1_2_1_38_1","first-page":"329","volume-title":"Proceedings of IEEE International Conference on Software Maintenance and Evolution, ICSME 2023","author":"Rodr\u00edguez-C\u00e1rdenas Daniel","year":"2023","unstructured":"Daniel Rodr\u00edguez-C\u00e1rdenas, David N. Palacio, Dipin Khati, Henry Burke, and Denys Poshyvanyk. 2023. Benchmarking Causal Study to Interpret Large Language Models for Source Code. In Proceedings of IEEE International Conference on Software Maintenance and Evolution, ICSME 2023, Bogot\u00e1, Colombia, October 1-6, 2023. IEEE, 329-334."},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems, NeurIPS 2023","author":"Rosa Biagio La","year":"2023","unstructured":"Biagio La Rosa, Leilani Gilpin, and Roberto Capobianco. 2023. Towards a fuller understanding of neurons with Clustered Compositional Explanations. In Proceedings of Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems, NeurIPS 2023, New Orleans, LA, USA, December 10 -16, 2023."},{"key":"e_1_2_1_40_1","first-page":"107","volume-title":"Proceedings of the 2024 ACM\/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, NIER@ICSE 2024","author":"Saad Mootez","year":"2024","unstructured":"Mootez Saad and Tushar Sharma. 2024. Naturalness of Attention: Revisiting Attention in Code Language Models. In Proceedings of the 2024 ACM\/IEEE 44th International Conference on Software Engineering: New Ideas and Emerging Results, NIER@ICSE 2024, Lisbon, Portugal, April 14-20, 2024. ACM, 107-111."},{"key":"e_1_2_1_41_1","first-page":"24","volume-title":"Proceedings of the 39th IEEE\/ACM International Conference on Automated Software Engineering Workshops, ASEW 2024","author":"Sharma Arushi","year":"2024","unstructured":"Arushi Sharma, Zefu Hu, Christopher J. Quinn, and Ali Jannesari. 2024. Redundancy and Concept Analysis for Code-trained Language Models. In Proceedings of the 39th IEEE\/ACM International Conference on Automated Software Engineering Workshops, ASEW 2024, Sacramento, CA, USA, 27 October 2024 -1 November 2024. ACM, 24-34."},{"key":"e_1_2_1_42_1","first-page":"437","volume-title":"Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension, ICPC 2022","author":"Sharma Rishab","year":"2022","unstructured":"Rishab Sharma, Fuxiang Chen, Fatemeh H. Fard, and David Lo. 2022. An exploratory study on code attention in BERT. In Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension, ICPC 2022, Virtual Event, May 16-17, 2022. ACM, 437-448."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track Proceedings.","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. In Proceedings of 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track Proceedings."},{"key":"e_1_2_1_44_1","first-page":"1882","volume-title":"Proceedings of 47th IEEE\/ACM International Conference on Software Engineering, ICSE 2025","author":"Sun Weisong","year":"2025","unstructured":"Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang, Chunrong Fang, Yi Liu, Gelei Deng, Yang Liu, and Zhenyu Chen. 2025. Source Code Summarization in the Era of Large Language Models. In Proceedings of 47th IEEE\/ACM International Conference on Software Engineering, ICSE 2025, Ottawa, ON, Canada, April 26 -May 6, 2025. IEEE, 1882-1894."},{"key":"e_1_2_1_45_1","first-page":"3319","volume-title":"Proceedings of the 34th International Conference on Machine Learning, ICML 2017","author":"Sundararajan Mukund","year":"2017","unstructured":"Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research, Vol. 70). PMLR, 3319-3328."},{"key":"e_1_2_1_46_1","first-page":"5701","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024","author":"Tang Tianyi","year":"2024","unstructured":"Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024. Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024. Association for Computational Linguistics, 5701-5715."},{"key":"e_1_2_1_47_1","first-page":"413","volume-title":"Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension, ICPC 2022","author":"Tao Chenning","year":"2022","unstructured":"Chenning Tao, Qi Zhan, Xing Hu, and Xin Xia. 2022. C4: contrastive cross-language code clone detection. In Proceedings of the 30th IEEE\/ACM International Conference on Program Comprehension, ICPC 2022, Virtual Event, May 16-17, 2022. ACM, 413-424."},{"key":"e_1_2_1_48_1","series-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL","volume-title":"Long Papers","author":"Tenney Ian","year":"2019","unstructured":"Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT Rediscovers the Classical NLP Pipeline. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28-August 2, 2019, Volume 1: Long Papers. Association for Computational Linguistics, 4593-4601."},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems, NeurIPS 2020","author":"Tran Chau","year":"2020","unstructured":"Chau Tran, Yuqing Tang, Xian Li, and Jiatao Gu. 2020. Cross-lingual Retrieval for Iterative Self-Supervised Training. In Proceedings of Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems, NeurIPS 2020, December 6-12, 2020, virtual."},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2022","author":"Troshin Sergey","year":"2022","unstructured":"Sergey Troshin and Nadezhda Chirkova. 2022. Probing Pretrained Models of Source Codes. In Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2022, Abu Dhabi, United Arab Emirates (Hybrid), December 8, 2022. Association for Computational Linguistics, 371-383."},{"key":"e_1_2_1_51_1","first-page":"1823","volume-title":"Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM 2019","author":"Winter Benjamin","year":"2019","unstructured":"van Aken, Benjamin Winter, Alexander L\u00f6ser, and Felix A. Gers. 2019. How Does BERT Answer Questions?: A Layer-Wise Analysis of Transformer Representations. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM 2019, Beijing, China, November 3-7, 2019. ACM, 1823-1832."},{"key":"e_1_2_1_52_1","first-page":"2377","volume-title":"Proceedings of IEEE\/ACM 44th International Conference on Software Engineering, ICSE 2022","author":"Wan Yao","year":"2022","unstructured":"Yao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui, Guandong Xu, and Hai Jin. 2022. What Do They Capture? -A Structural Analysis of Pre-Trained Language Models for Source Code. In Proceedings of IEEE\/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022. ACM, 2377-2388."},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of IEEE\/ACM 44th International Conference on Software Engineering, ICSE 2026 SEIP. Rio de Janeiro","author":"Wang Chaofan","year":"2026","unstructured":"Chaofan Wang, Tingrui Yu, Beijun Shen, Jie Wang, Dong Chen, Wenrui Zhang, Yuling Shi, Chen Xie, and Xiaodong Gu. 2026. EvoC2Rust: A Skeleton-guided Framework for Project-Level C-to-Rust Translation. In Proceedings of IEEE\/ACM 44th International Conference on Software Engineering, ICSE 2026 SEIP. Rio de Janeiro, Brazil, April 12-18, 2026. ACM."},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the Eleventh International Conference on Learning Representations, ICLR 2023","author":"Wang Kevin Ro","year":"2023","unstructured":"Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small. In Proceedings of the Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023."},{"key":"e_1_2_1_55_1","first-page":"801","volume-title":"Proc. ACM Program. Lang. 7, OOPSLA2","author":"Wang Yu","year":"2023","unstructured":"Yu Wang, Ke Wang, and Linzhang Wang. 2023. An Explanation Method for Models of Code. Proc. ACM Program. Lang. 7, OOPSLA2 (2023), 801-827."},{"key":"e_1_2_1_56_1","volume-title":"Opportunities, and Challenges of Representation Engineering for Large Language Models. CoRR abs\/2502.19649","author":"Wehner Jan","year":"2025","unstructured":"Jan Wehner, Sahar Abdelnabi, Daniel Tan, David Krueger, and Mario Fritz. 2025. Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models. CoRR abs\/2502.19649 (2025)."},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 -Demos","author":"Wolf Thomas","year":"2020","unstructured":"Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R\u00e9mi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, EMNLP 2020 -Demos, Online, November 16-20, 2020. Association for Computational Linguistics, 38-45."},{"key":"e_1_2_1_58_1","first-page":"9393","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE","author":"Xu Haoyun","year":"2025","unstructured":"Haoyun Xu, Runzhe Zhan, Yingpeng Ma, Derek F. Wong, and Lidia S. Chao. 2025. Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model. In Proceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19-24, 2025. Association for Computational Linguistics, 9393-9406."},{"key":"e_1_2_1_59_1","volume-title":"What does Transformer learn about source code? CoRR abs\/2207.08466","author":"Zhang Kechi","year":"2022","unstructured":"Kechi Zhang, Ge Li, and Zhi Jin. 2022. What does Transformer learn about source code? CoRR abs\/2207.08466 (2022)."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639372"},{"key":"e_1_2_1_61_1","doi-asserted-by":"crossref","first-page":"5673","DOI":"10.1145\/3580305.3599790","volume-title":"Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2023","author":"Zheng Qinkai","year":"2023","unstructured":"Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2023, Long Beach, CA, USA, August 6-10, 2023. ACM, 5673-5684."},{"key":"e_1_2_1_62_1","unstructured":"Andy Zou Long Phan Sarah Li Chen James Campbell Phillip Guo Richard Ren Alexander Pan Xuwang Yin Mantas Mazeika Ann-Kathrin Dombrowski Shashwat Goel Nathaniel Li Michael J. Byun Zifan Wang Alex Mallen Steven Basart Sanmi Koyejo Dawn Song Matt Fredrikson J. Zico Kolter and Dan Hendrycks. 2023. Representation Engineering: A Top-Down Approach to AI Transparency. CoRR abs\/2310.01405 (2023)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797083","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:58:59Z","timestamp":1782842339000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3797083"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":62,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797083"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3797083","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}