{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:11:35Z","timestamp":1782846695137,"version":"3.54.5"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Code editing constitutes a fundamental practice in software development, wherein developers modify existing codebases according to natural language requirements. Accurate code editing necessitates a comprehensive understanding of both the existing codebase and the modification requirements. Although large language models (LLMs) have demonstrated promising performance in code editing tasks, they suffer from substantial inefficiency by generating entire modified files that largely consist of unchanged code. While smaller models could potentially address this inefficiency, they typically lack the capacity to effectively comprehend long code contexts required for accurate editing. To ensure both effectiveness and efficiency, we propose to decompose code editing into a two-stage cascade:\n                    <jats:bold>edit sketch generation<\/jats:bold>\n                    , wherein a large model first produces concise sketches representing the requisite modifications (the more challenging phase), and\n                    <jats:bold>edit sketch application<\/jats:bold>\n                    , wherein a smaller model integrates these sketches into the original code to produce the final output edited code (the simpler phase). This cascaded design reduces the number of tokens generated by the large model, as the majority of the output is handled by the smaller, more efficient model, thereby enhancing overall efficiency. However, the effectiveness of this approach is constrained by current small models\u2019 limited capabilities in handling long-context scenarios and cross-file dependencies, which are essential for accurate sketch application in real-world codebases. To address these limitations and enhance smaller models\u2019 sketch application capabilities, we introduce the first large-scale sketch application dataset comprising over 100K training instances and 800M tokens, along with a human-evaluated benchmark, and propose specialized training strategies including curriculum-based long-context training and multi-file augmentation. Our comprehensive experiments demonstrate that our cascaded framework inherently reduces inference costs compared to direct editing with large models. Furthermore, combining large models with our fine-tuned smaller models can achieve even superior performance. For instance, on the Aider benchmark, employing DeepSeek R1 as the edit sketch generation model alongside a fine-tuned Qwen2.5 Coder 14B model for the application phase improves Pass@2 11.1% compared to direct editing with DeepSeek R1 alone. Additionally, the cascaded approach reduces execution time and cost by 13% and 19%, respectively, demonstrating both performance gains and efficiency improvements.\n                  <\/jats:p>","DOI":"10.1145\/3808101","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"2094-2116","source":"Crossref","is-referenced-by-count":0,"title":["Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-3935-7328","authenticated-orcid":false,"given":"Chaozheng","family":"Wang","sequence":"first","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0008-9092-3381","authenticated-orcid":false,"given":"Zezhou","family":"Yang","sequence":"additional","affiliation":[{"name":"Hong Kong University, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-8102-480X","authenticated-orcid":false,"given":"Shuzheng","family":"Gao","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-8513-6836","authenticated-orcid":false,"given":"Cuiyun","family":"Gao","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-9897-4086","authenticated-orcid":false,"given":"Zongjie","family":"Li","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-8370-644X","authenticated-orcid":false,"given":"Yichen","family":"Li","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0003-6970-0857","authenticated-orcid":false,"given":"Ting","family":"Peng","sequence":"additional","affiliation":[{"name":"Tencent, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0004-0655-9398","authenticated-orcid":false,"given":"Hailiang","family":"Huang","sequence":"additional","affiliation":[{"name":"Tencent, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0003-7060-4109","authenticated-orcid":false,"given":"Yuetang","family":"Deng","sequence":"additional","affiliation":[{"name":"Tencent, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-3666-5798","authenticated-orcid":false,"given":"Michael R.","family":"Lyu","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Guangzhou, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Forty-second International Conference on Machine Learning. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=3B6fF1PxYD","author":"Aggarwal Tushar","year":"2025","unstructured":"Tushar Aggarwal, Swayam Singh, Abhijeet Awasthi, Aditya Kanade, and Nagarajan Natarajan. 2025. NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits. In Forty-second International Conference on Machine Learning. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=3B6fF1PxYD"},{"key":"e_1_2_1_2_1","unstructured":"Aider. 2025. Aider LLM Leaderboards. https:\/\/2.zoppoz.workers.dev:443\/https\/aider.chat\/docs\/leaderboards\/."},{"key":"e_1_2_1_3_1","unstructured":"Aider. 2025. Aider's polyglot benchmark. https:\/\/2.zoppoz.workers.dev:443\/https\/aider.chat\/2024\/12\/21\/polyglot.html#the-polyglot-benchmark."},{"key":"e_1_2_1_4_1","unstructured":"Anthropic. 2025. Introducing Claude 4. https:\/\/2.zoppoz.workers.dev:443\/https\/www.anthropic.com\/news\/claude-4."},{"key":"e_1_2_1_5_1","volume-title":"Enhancing llm code generation: A systematic evaluation of multi-agent collaboration and runtime debugging for improved accuracy, reliability, and latency. arXiv preprint arXiv:2505.02133","author":"Ashrafi Nazmus","year":"2025","unstructured":"Nazmus Ashrafi, Salah Bouktif, and Mohammed Mediani. 2025. Enhancing llm code generation: A systematic evaluation of multi-agent collaboration and runtime debugging for improved accuracy, reliability, and latency. arXiv preprint arXiv:2505.02133 (2025)."},{"key":"e_1_2_1_6_1","first-page":"675","volume-title":"Proc. ACM Softw. Eng. 1, FSE","author":"Bairi Ramakrishna","year":"2024","unstructured":"Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D. C., Arun Iyer, Suresh Parthasarathy, Sriram K. Rajamani, Balasubramanyan Ashok, and Shashank Shet. 2024. CodePlan: Repository-Level Coding using LLMs and Planning. Proc. ACM Softw. Eng. 1, FSE (2024), 675-698."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.44.4.433"},{"key":"e_1_2_1_8_1","volume-title":"First Conference on Language Modeling.","author":"Cassano Federico","year":"2024","unstructured":"Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Jacob Ginesin, Edward Berman, George Chakhnashvili, Anton Lozhkov, Carolyn Jane Anderson, et al. 2024. Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions. In First Conference on Language Modeling."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2020.3020502"},{"key":"e_1_2_1_10_1","first-page":"1755","volume-title":"Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming. In 47th IEEE\/ACM International Conference on Software Engineering, ICSE 2025","author":"Chakraborty Saikat","year":"2025","unstructured":"Saikat Chakraborty, Gabriel Ebner, Siddharth Bhat, Sarah Fakhoury, Sakina Fatima, Shuvendu K. Lahiri, and Nikhil Swamy. 2025. Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming. In 47th IEEE\/ACM International Conference on Software Engineering, ICSE 2025, Ottawa, ON, Canada, April 26 -May 6, 2025. IEEE, 1755-1767."},{"key":"e_1_2_1_11_1","first-page":"443","volume-title":"On Multi-Modal Learning of Editing Source Code. In 36th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2021","author":"Chakraborty Saikat","year":"2021","unstructured":"Saikat Chakraborty and Baishakhi Ray. 2021. On Multi-Modal Learning of Editing Source Code. In 36th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2021, Melbourne, Australia, November 15-19, 2021. IEEE, 443-455."},{"key":"e_1_2_1_12_1","volume-title":"Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174 (2016)."},{"key":"e_1_2_1_13_1","volume-title":"The Twelfth International Conference on Learning Representations, ICLR 2024","author":"Chen Xinyun","year":"2024","unstructured":"Xinyun Chen, Maxwell Lin, Nathanael Sch\u00e4rli, and Denny Zhou. 2024. Teaching Large Language Models to Self- Debug. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net."},{"key":"e_1_2_1_14_1","first-page":"572","volume-title":"ChatUniTest: A Framework for LLM-Based Test Generation. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, Porto de Galinhas","author":"Chen Yinghao","year":"2024","unstructured":"Yinghao Chen, Zehao Hu, Chen Zhi, Junxiao Han, Shuiguang Deng, and Jianwei Yin. 2024. ChatUniTest: A Framework for LLM-Based Test Generation. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, Porto de Galinhas, Brazil, July 15-19, 2024, Marcelo d'Amorim (Ed.). ACM, 572-576."},{"key":"e_1_2_1_15_1","unstructured":"Clang LLVM. 2025. ClangFormat. https:\/\/2.zoppoz.workers.dev:443\/https\/clang.llvm.org\/docs\/ClangFormat.html."},{"key":"e_1_2_1_16_1","unstructured":"CurSor. 2025. The AI Code Editor. https:\/\/2.zoppoz.workers.dev:443\/https\/cursor.com\/en."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1189"},{"key":"e_1_2_1_18_1","unstructured":"DeepInfra. 2025. Simple Pricing Deep Infrastructure. https:\/\/2.zoppoz.workers.dev:443\/https\/deepinfra.com\/pricing."},{"key":"e_1_2_1_19_1","unstructured":"Google. 2025. Gemini 2.5 Pro. https:\/\/2.zoppoz.workers.dev:443\/https\/deepmind.google\/models\/gemini\/pro\/."},{"key":"e_1_2_1_20_1","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)."},{"key":"e_1_2_1_21_1","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y Wu YK Li et al. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming-The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196 (2024)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","unstructured":"Siming Huang Tianhao Cheng Jason Klein Liu Weidi Xu Jiaran Hao Liuyihan Song Yang Xu Jian Yang Jiaheng Liu Chenchen Zhang Linzheng Chai Ruifeng Yuan Xianzhen Luo Qiufeng Wang YuanTao Fan Qingfu Zhu Zhaoxiang Zhang Yang Gao Jie Fu Qian Liu Houyi Li Ge Zhang Yuan Qi Xu Yinghui Wei Chu and Zili Wang. 2025. OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) Wanxiang Che Joyce Nabende Ekaterina Shutova and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics Vienna Austria 33167-33193. doi:10.18653\/v1\/2025.acl-long.1591 10.18653\/v1\/2025.acl-long.1591","DOI":"10.18653\/v1\/2025.acl-long.1591"},{"key":"e_1_2_1_23_1","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu et al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186 (2024)."},{"key":"e_1_2_1_24_1","volume-title":"Decoupled Weight Decay Regularization. International Conference on Learning Representations, ICLR","author":"Ilya Loshchilov","year":"2018","unstructured":"Loshchilov Ilya and Hutter Frank. 2018. Decoupled Weight Decay Regularization. International Conference on Learning Representations, ICLR (2018)."},{"key":"e_1_2_1_25_1","volume-title":"Large language models (llms) for source code analysis: applications, models and datasets. arXiv preprint arXiv:2503.17502","author":"Jelodar Hamed","year":"2025","unstructured":"Hamed Jelodar, Mohammad Meymani, and Roozbeh Razavi-Far. 2025. Large language models (llms) for source code analysis: applications, models and datasets. arXiv preprint arXiv:2503.17502 (2025)."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, CCS 2025","author":"Ji Zimo","year":"2025","unstructured":"Zimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma, Zongjie Li, and Shuai Wang. 2025. Measuring and Augmenting Large Language Models for Solving Offensive Security Challenges. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, CCS 2025, Taipei, Taiwan, October 13-17, 2025."},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","first-page":"6","DOI":"10.5120\/20707-3021","article-title":"A review on software maintenance issues and how to reduce maintenance efforts","volume":"118","author":"Kaur Uttamjit","year":"2015","unstructured":"Uttamjit Kaur and Gagandeep Singh. 2015. A review on software maintenance issues and how to reduce maintenance efforts. International Journal of Computer Applications 118, 1 (2015), 6-11.","journal-title":"International Journal of Computer Applications"},{"key":"e_1_2_1_28_1","unstructured":"Tobias Kuipers. 2016. Why you need to know about code maintainability. https:\/\/2.zoppoz.workers.dev:443\/https\/www.oreilly.com\/content\/why-youneed-to-know-about-code-maintainability\/."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_2_1_30_1","article-title":"CodeEditor: Learning to Edit Source Code with Pre-trained Models","volume":"32","author":"Li Jia","year":"2023","unstructured":"Jia Li, Ge Li, Zhuo Li, Zhi Jin, Xing Hu, Kechi Zhang, and Zhiyi Fu. 2023. CodeEditor: Learning to Edit Source Code with Pre-trained Models. ACM Trans. Softw. Eng. Methodol. 32, 6 (2023), 143:1-143:22.","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"e_1_2_1_31_1","volume-title":"Hui Chen, Yuxi Xie, Tiedong Liu, Michael Shieh, and Junxian He.","author":"Li Kaixin","year":"2024","unstructured":"Kaixin Li, Qisheng Hu, James Xu Zhao, Hui Chen, Yuxi Xie, Tiedong Liu, Michael Shieh, and Junxian He. 2024. InstructCoder: Instruction Tuning Large Language Models for Code Editing, Xiyan Fu and Eve Fleisig (Eds.)."},{"key":"e_1_2_1_32_1","first-page":"1238","volume-title":"CCTEST: Testing and Repairing Code Completion Systems. In 45th IEEE\/ACM International Conference on Software Engineering, ICSE 2023","author":"Li Zongjie","year":"2023","unstructured":"Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023. CCTEST: Testing and Repairing Code Completion Systems. In 45th IEEE\/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 2023. IEEE, 1238-1250."},{"key":"e_1_2_1_33_1","first-page":"786","volume-title":"Proceedings of the ACM on Programming Languages 9, OOPSLA1","author":"Li Zongjie","year":"2025","unstructured":"Zongjie Li, Daoyuan Wu, Shuai Wang, and Zhendong Su. 2025. Api-guided dataset synthesis to finetune large code models. Proceedings of the ACM on Programming Languages 9, OOPSLA1 (2025), 786-815."},{"key":"e_1_2_1_34_1","unstructured":"Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)."},{"key":"e_1_2_1_35_1","first-page":"466","volume-title":"Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024","author":"Liu Chenyan","year":"2024","unstructured":"Chenyan Liu, Yufan Cai, Yun Lin, Yuhuan Huang, Yunrui Pei, Bo Jiang, Ping Yang, Jin Song Dong, and Hong Mei. 2024. CoEdPilot: Recommending Code Edits with Learned Prior Edit Relevance, Project-wise Awareness, and Interactive Nature. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2024, Vienna, Austria, September 16-20, 2024. ACM, 466-478."},{"key":"e_1_2_1_36_1","volume-title":"SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations.","author":"Loshchilov Ilya","year":"2016","unstructured":"Ilya Loshchilov and Frank Hutter. 2016. SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations."},{"key":"e_1_2_1_37_1","volume-title":"Reusing software: Issues and research directions","author":"Mili Hafedh","year":"2002","unstructured":"Hafedh Mili, Fatma Mili, and Ali Mili. 2002. Reusing software: Issues and research directions. IEEE transactions on Software Engineering 21, 6 (2002), 528-562."},{"key":"e_1_2_1_38_1","volume-title":"NeurIPS 2023 workshop on instruction tuning and instruction following.","author":"Muennighoff Niklas","year":"2023","unstructured":"Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro Von Werra, and Shayne Longpre. 2023. Octopack: Instruction tuning code large language models. In NeurIPS 2023 workshop on instruction tuning and instruction following."},{"key":"e_1_2_1_39_1","volume-title":"Prompting LLMs for Code Editing: Struggles and Remedies. CoRR abs\/2504.20196","author":"Nam Daye","year":"2025","unstructured":"Daye Nam, Ahmed Omran, Ambar Murillo, Saksham Thakur, Abner Araujo, Marcel Blistein, Alexander Fr\u00f6mmgen, Vincent J. Hellendoorn, and Satish Chandra. 2025. Prompting LLMs for Code Editing: Struggles and Remedies. CoRR abs\/2504.20196 (2025)."},{"key":"e_1_2_1_40_1","unstructured":"OpenAI. 2025. Introducing gpt-oss. https:\/\/2.zoppoz.workers.dev:443\/https\/openai.com\/index\/introducing-gpt-oss\/."},{"key":"e_1_2_1_41_1","unstructured":"Qwen3. 2025. Qwen3 Coder -Agentic Coding Adventure. https:\/\/2.zoppoz.workers.dev:443\/https\/qwen3lm.com\/."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00024"},{"key":"e_1_2_1_43_1","first-page":"551","volume-title":"2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Ren Jie","year":"2021","unstructured":"Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. {Zero-offload}: Democratizing {billion-scale} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). 551-564."},{"key":"e_1_2_1_44_1","volume-title":"Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297","author":"Ren Shuo","year":"2020","unstructured":"Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020. Codebleu: a method for automatic evaluation of code synthesis. arXiv preprint arXiv:2009.10297 (2020)."},{"key":"e_1_2_1_45_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al.","author":"Rozi\u00e8re Baptiste","year":"2023","unstructured":"Baptiste Rozi\u00e8re, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_2_1_46_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al.","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_2_1_47_1","volume-title":"Simhash: Hash-based similarity detection. Technical Report. Technical report, Google.","author":"Sadowski Caitlin","year":"2007","unstructured":"Caitlin Sadowski and Greg Levin. 2007. Simhash: Hash-based similarity detection. Technical Report. Technical report, Google."},{"key":"e_1_2_1_48_1","unstructured":"ByteDance Seed Yuyu Zhang Jing Su Yifan Sun Chenguang Xi Xia Xiao Shen Zheng Anxiang Zhang Kaibo Liu Daoguang Zan et al. 2025. Seed-Coder: Let the Code Model Curate Data for Itself. arXiv preprint arXiv:2506.03524 (2025)."},{"key":"e_1_2_1_49_1","unstructured":"Yuxuan Song Zheng Zhang Cheng Luo Pengyang Gao Fan Xia Hao Luo Zheng Li Yuehang Yang Hongli Yu Xingwei Qu et al. 2025. Seed diffusion: A large-scale diffusion language model with high-speed inference. arXiv preprint arXiv:2508.02193 (2025)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2393596.2393656"},{"key":"e_1_2_1_51_1","first-page":"1567","volume-title":"Proceedings of the ACM on Software Engineering 2, FSE","author":"Wang Chaozheng","year":"2025","unstructured":"Chaozheng Wang, Jia Feng, Shuzheng Gao, Cuiyun Gao, Zongjie Li, Ting Peng, Hailiang Huang, Yuetang Deng, and Michael Lyu. 2025. Beyond PEFT: Layer-Wise Optimization for More Effective and Efficient Large Code Model Tuning. Proceedings of the ACM on Software Engineering 2, FSE (2025), 1567-1590."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695004"},{"key":"e_1_2_1_53_1","volume-title":"Exploring Multi-Lingual Bias of Large Code Models in Code Generation. arXiv preprint arXiv:2404.19368","author":"Wang Chaozheng","year":"2024","unstructured":"Chaozheng Wang, Zongjie Li, Cuiyun Gao, Wenxuan Wang, Ting Peng, Hailiang Huang, Yuetang Deng, Shuai Wang, and Michael R Lyu. 2024. Exploring Multi-Lingual Bias of Large Code Models in Code Generation. arXiv preprint arXiv:2404.19368 (2024)."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3696630.3728535"},{"key":"e_1_2_1_55_1","first-page":"690","volume-title":"Proceedings of the ACM on Software Engineering 2, FSE","author":"Wang Yanlin","year":"2025","unstructured":"Yanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen, Mingzhi Mao, Xilin Liu, Yuchi Ma, and Zibin Zheng. 2025. Beyond functional correctness: Investigating coding style inconsistencies in large language models. Proceedings of the ACM on Software Engineering 2, FSE (2025), 690-712."},{"key":"e_1_2_1_56_1","volume-title":"Magicoder: Source code is all you need. arXiv preprint arXiv:2312.02120","author":"Wei Yuxiang","year":"2023","unstructured":"Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2023. Magicoder: Source code is all you need. arXiv preprint arXiv:2312.02120 (2023)."},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the 34th ACM SIGSOFT International Symposium on Software Testing and Analysis.","author":"Wong Wai Kin","year":"2025","unstructured":"Wai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li, Zhibo Liu, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu. 2025. DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled Code. In Proceedings of the 34th ACM SIGSOFT International Symposium on Software Testing and Analysis."},{"key":"e_1_2_1_58_1","unstructured":"An Yang Anfeng Li Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chang Gao Chengen Huang Chenxu Lv et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025)."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.280"},{"key":"e_1_2_1_60_1","volume-title":"Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input Generators. arXiv preprint arXiv:2501.19282","author":"Zhang Kunpeng","year":"2025","unstructured":"Kunpeng Zhang, Zongjie Li, Daoyuan Wu, Shuai Wang, and Xin Xia. 2025. Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input Generators. arXiv preprint arXiv:2501.19282 (2025)."},{"key":"e_1_2_1_61_1","volume-title":"Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al.","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al. 2024. Sglang: Efficient execution of structured language model programs. Advances in neural information processing systems 37 (2024), 62557-62583."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3808101","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:34:12Z","timestamp":1782844452000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3808101"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":61,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808101"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3808101","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}