{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:59:28Z","timestamp":1782845968193,"version":"3.54.5"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2129824, 2312487, 2217071"],"award-info":[{"award-number":["2129824, 2312487, 2217071"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>With the increasing popularity of large language models (LLMs) and LLM-based agents, reliable and effective code evaluation metrics (CEMs) have become crucial for progress across several software engineering tasks. While popular benchmarks often provide test cases to assess the correctness of generated code, crafting and executing test cases is expensive. Reference-based CEMs provide a cheaper alternative by scoring a candidate program based on its functional similarity to a reference. Although prior research has focused on reporting the weak correlation between these CEMs and functional correctness, the causes are only assumed, and plausible solutions remain unexplored. In this work, we critically evaluate four state-of-the-art reference-based CEMs, revealing their strong bias towards surface-level features rather than code functionality. Despite this  \nsurface bias, current evaluation datasets for these CEMs rarely include code pairs that are surface-similar yet functionally dissimilar, or functionally similar yet surface-dissimilar.<\/jats:p>\n                  <jats:p>To mitigate this gap, we propose LoCaL (Looks Can Lie), a CEM evaluation benchmark, with 3117 code  \npairs at both the method and program levels. Each pair is labeled with a functional similarity score and aims to target regions where CEMs are likely to perform poorly. The functional similarity scores are calculated through differential fuzzing, which eliminates the need for predefined test cases and, at the same time, improves the reliability of the scores by executing an order of magnitude more tests than prior work. We find that all four CEMs show significant performance degradation on LoCaL, compared to the baselines. Finally, based on our findings, we draw the implication that exposing CEMs to LoCaL-like data might facilitate the development of metrics that are robust to surface bias.<\/jats:p>","DOI":"10.1145\/3797089","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"1357-1380","source":"Crossref","is-referenced-by-count":0,"title":["LoCaL: Countering Surface Bias in Code Evaluation Metrics"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-5511-9307","authenticated-orcid":false,"given":"Simantika Bhattacharjee","family":"Dristi","sequence":"first","affiliation":[{"name":"University of Virginia, Charlottesville, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-1937-1544","authenticated-orcid":false,"given":"Matthew B.","family":"Dwyer","sequence":"additional","affiliation":[{"name":"University of Virginia, Charlottesville, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"d.]. ShareCode. https:\/\/2.zoppoz.workers.dev:443\/https\/sharecode.io\/. Online programming competition platform","year":"2025","unstructured":"n. d.]. ShareCode. https:\/\/2.zoppoz.workers.dev:443\/https\/sharecode.io\/. Online programming competition platform; accessed 2025-08-11."},{"key":"e_1_2_1_2_1","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le and Charles Sutton. 2021. Program Synthesis with Large Language Models. arXiv:2108.07732 [cs.PL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_2_1_3_1","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization Jade Goldstein Alon Lavie Chin-Yew Lin and Clare Voss (Eds.). Association for Computational Linguistics Ann Arbor Michigan 65-72. https:\/\/2.zoppoz.workers.dev:443\/https\/aclanthology.org\/W05-0909\/"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3586030"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSM"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2009.70"},{"key":"e_1_2_1_7_1","unstructured":"Bei Chen Fengji Zhang Anh Nguyen Daoguang Zan Zeqi Lin Jian-Guang Lou and Weizhu Chen. 2022. CodeT: Code Generation with Generated Tests. arXiv:2207.10397 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2207.10397"},{"key":"e_1_2_1_8_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 [cs.LG] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3423769"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2022. FLASHATTENTION: fast and memoryefficient exact attention with IO-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '22). Curran Associates Inc., Red Hook, NY, USA, Article 1189, 16 pages."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3695991"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","unstructured":"Yihong Dong Ge Li Xue Jiang and Zhi Jin. 2023. Antecedent Predictions Are More Important Than You Think: An Effective Method for Tree-Based Code Generation. doi:10.3233\/FAIA230317 10.3233\/FAIA230317","DOI":"10.3233\/FAIA230317"},{"key":"e_1_2_1_13_1","volume-title":"Bin Ji, Qian Liu, and See-Kiong Ng.","author":"Du Mingzhe","year":"2024","unstructured":"Mingzhe Du, Anh Tuan Luu, Bin Ji, Qian Liu, and See-Kiong Ng. 2024. Mercury: A Code Efficiency Benchmark for Code Large Language Models. arXiv:2402.07844 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2402.07844"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556903"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3472476"},{"key":"e_1_2_1_17_1","volume-title":"Atheris: A Coverage-Guided Python Fuzzer. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/google\/atheris Accessed: 2025-08-12.","year":"2020","unstructured":"Google. 2020. Atheris: A Coverage-Guided Python Fuzzer. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/google\/atheris Accessed: 2025-08-12."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3236024.3264835"},{"key":"e_1_2_1_19_1","unstructured":"Konrad Halas. 2013. MutPy: Mutation Testing Tool for Python 3.x. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/mutpy\/mutpy Accessed: 2025-08-12."},{"key":"e_1_2_1_20_1","volume-title":"Ahmad Humayun, Waris Gill, Abdul Haddi Amjad, Ali R. Butt, Mohammad Taha Khan, and Muhammad Ali Gulzar.","author":"Haroon Sabaat","year":"2025","unstructured":"Sabaat Haroon, Ahmad Faraz Khan, Ahmad Humayun, Waris Gill, Abdul Haddi Amjad, Ali R. Butt, Mohammad Taha Khan, and Muhammad Ali Gulzar. 2025. How Accurately Do Large Language Models Understand Code? arXiv:2504.04372 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2504.04372"},{"key":"e_1_2_1_21_1","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=sD93GOzH3i5"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.269"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3747588"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/30.1-2.81"},{"key":"e_1_2_1_25_1","volume-title":"SPoC: search-based pseudocode to code","author":"Kulal Sumith","unstructured":"Sumith Kulal, Panupong Pasupat, Kartik Chandra, Mina Lee, Oded Padon, Alex Aiken, and Percy Liang. 2019. SPoC: search-based pseudocode to code. Curran Associates Inc., Red Hook, NY, USA."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","unstructured":"Jierui Li Hung L\u00ea Yinbo Zhou Caiming Xiong Silvio Savarese and Doyen Sahoo. 2024. CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models. doi:10.48550\/arXiv.2411.04329 10.48550\/arXiv.2411.04329","DOI":"10.48550\/arXiv.2411.04329"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","unstructured":"Yujia Li David Choi Junyoung Chung Nate Kushman Julian Schrittwieser R\u00e9mi Leblond Tom Eccles James Keeling Felix Gimeno Agustin Dal Lago Thomas Hubert Peter Choy Cyprien de Masson d'Autume Igor Babuschkin Xinyun Chen Po-Sen Huang Johannes Welbl Sven Gowal Alexey Cherepanov James Molloy Daniel J. Mankowitz Esme Sutherland Robson Pushmeet Kohli Nando de Freitas Koray Kavukcuoglu and Oriol Vinyals. 2022. Competition-level code generation with AlphaCode. Science 378 6624 (2022) 1092-1097. arXiv:https:\/\/2.zoppoz.workers.dev:443\/https\/www.science.org\/doi\/pdf\/10.1126\/science.abq1158 doi:10.1126\/science.abq1158 10.1126\/science.abq1158","DOI":"10.1126\/science.abq1158"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.120073"},{"key":"e_1_2_1_29_1","volume-title":"ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74-81. https:\/\/2.zoppoz.workers.dev:443\/https\/aclanthology.org\/W04-1013\/"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/QRS60937.2023.00066"},{"key":"e_1_2_1_31_1","unstructured":"Olle Lindgren. 2024. pyrefact: Automatic Python Refactoring. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/OlleLindgren\/pyrefact. GitHub repository accessed 2025-08-22."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3628159"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0943"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","unstructured":"Nickil Maveli Antonio Vergari and Shay Cohen. 2024. What can Large Language Models Capture about Code Functional Equivalence? doi:10.48550\/arXiv.2408.11081 10.48550\/arXiv.2408.11081","DOI":"10.48550\/arXiv.2408.11081"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.1976.233837"},{"key":"e_1_2_1_36_1","first-page":"100","article-title":"Differential testing for software","volume":"10","author":"McKeeman William M","year":"1998","unstructured":"William M McKeeman. 1998. Differential testing for software. Digital Technical Journal 10, 1 (1998), 100-107.","journal-title":"Digital Technical Journal"},{"key":"e_1_2_1_37_1","first-page":"19","article-title":"Memo","volume":"218","author":"Michie Donald","year":"1968","unstructured":"Donald Michie. 1968. \"Memo\" Functions and Machine Learning. Nature 218 (1968), 19-22. https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar. org\/CorpusID:4265138","journal-title":"Functions and Machine Learning. Nature"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3660810"},{"key":"e_1_2_1_39_1","unstructured":"Atharva Naik. 2024. On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation. arXiv:2405.01580 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2405.01580"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3241537"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s002360050048"},{"key":"e_1_2_1_42_1","unstructured":"OpenAI. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_2_1_43_1","unstructured":"OpenAI. 2025. ChatGPT. https:\/\/2.zoppoz.workers.dev:443\/https\/chat.openai.com. Large language model accessed 2025-09-11\"."},{"key":"e_1_2_1_44_1","unstructured":"OpenRewrite Project. 2025. rewrite-python: Automated Python Refactoring Using Rewrite. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/ openrewrite\/rewrite-python. GitHub repository accessed 2025-08-22."},{"key":"e_1_2_1_45_1","unstructured":"Zhenyu Pan Rongyu Cao Yongchang Cao Yingwei Ma Binhua Li Fei Huang Han Liu and Yongbin Li. 2024. Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion? arXiv:2410.01353 [cs.SE] https: \/\/arxiv.org\/abs\/2410.01353"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/AITest62860.2024.00019"},{"key":"e_1_2_1_48_1","first-page":"253","article-title":"Mathematical Contributions to the Theory of Evolution","volume":"187","author":"Pearson Karl","year":"1896","unstructured":"Karl Pearson. 1896. Mathematical Contributions to the Theory of Evolution. III. Regression, Heredity, and Panmixia. Philosophical Transactions of the Royal Society of London. Series A 187 (1896), 253-318.","journal-title":"III. Regression, Heredity, and Panmixia. Philosophical Transactions of the Royal Society of London. Series A"},{"key":"e_1_2_1_49_1","volume-title":"COFFE: A Code Efficiency Benchmark for Code Generation. arXiv:2502.02827 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2502.02827","author":"Peng Yun","year":"2025","unstructured":"Yun Peng, Jun Wan, Yichen Li, and Xiaoxue Ren. 2025. COFFE: A Code Efficiency Benchmark for Code Generation. arXiv:2502.02827 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2502.02827"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-3049"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.213"},{"key":"e_1_2_1_52_1","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. CodeBLEU: a Method for Automatic Evaluation of Code Synthesis. arXiv:2009.10297 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2009.10297"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPC.2008.41"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/52.3-4.591"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575727"},{"key":"e_1_2_1_56_1","volume-title":"Learning Performance-Improving Code Edits. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id= ix7rLVHXyY","author":"Shypula Alexander G","year":"2024","unstructured":"Alexander G Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh. 2024. Learning Performance-Improving Code Edits. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id= ix7rLVHXyY"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/SCAM63643.2024.00028"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.19672949"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.2307\/1412159"},{"key":"e_1_2_1_60_1","volume-title":"Are NLP Metrics Suitable for Evaluating Generated Code?","author":"Takaichi Riku","unstructured":"Riku Takaichi, Yoshiki Higo, Shinsuke Matsumoto, Shinji Kusumoto, Toshiyuki Kurabayashi, Hiroyuki Kirinuki, and Haruto Tanno. 2022. Are NLP Metrics Suitable for Evaluating Generated Code?. In Product-Focused Software Process Improvement, Davide Taibi, Marco Kuhrmann, Tommi Mikkonen, Jil Kl\u00fcnder, and Pekka Abrahamsson (Eds.). Springer International Publishing, Cham, 531-537."},{"key":"e_1_2_1_61_1","doi-asserted-by":"crossref","unstructured":"Weixi Tong and Tianyi Zhang. 2024. CodeJudge: Evaluating Code Generation with Large Language Models. arXiv:2410.02184 [cs.LG] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2410.02184","DOI":"10.18653\/v1\/2024.emnlp-main.1118"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.3390\/fi16060188"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/icpc"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491101.3519665"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-"},{"key":"e_1_2_1_66_1","unstructured":"Guang Yang Yu Zhou Xiang Chen Wei Zheng Xing Hu Xin Zhou David Lo and Taolue Chen. 2025. CODE- DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation. arXiv:2505.19502 [cs.SE] https: \/\/arxiv.org\/abs\/2505.19502"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2002"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695537"},{"key":"e_1_2_1_69_1","volume-title":"CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code. In The 2023 Conference on Empirical Methods in Natural Language Processing. https: \/\/openreview.net\/forum?id=7cXoueVCoL","author":"Zhou Shuyan","year":"2023","unstructured":"Shuyan Zhou, Uri Alon, Sumit Agarwal, and Graham Neubig. 2023. CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code. In The 2023 Conference on Empirical Methods in Natural Language Processing. https: \/\/openreview.net\/forum?id=7cXoueVCoL"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21434"},{"key":"e_1_2_1_71_1","doi-asserted-by":"crossref","first-page":"2232","DOI":"10.18653\/v1\/2024.findings-eacl.148","volume-title":"Findings of the Association for Computational Linguistics: EACL 2024","author":"Zhuo Terry Yue","year":"2024","unstructured":"Terry Yue Zhuo. 2024. ICE-Score: Instructing Large Language Models to Evaluate Code. In Findings of the Association for Computational Linguistics: EACL 2024, Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian's, Malta, 2232-2242. https:\/\/2.zoppoz.workers.dev:443\/https\/aclanthology.org\/2024.findings-eacl.148\/"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797089","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797089","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:01:39Z","timestamp":1782842499000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3797089"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":71,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797089"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3797089","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}