{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:49:28Z","timestamp":1782848968675,"version":"3.54.5"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"PhD Scholarship Programme of Vingroup Innovation Foundation (VINIF), VinUniversity","award":["VINIF.2025.TS68"],"award-info":[{"award-number":["VINIF.2025.TS68"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Large language models for code (CodeLLMs) have demonstrated remarkable success in standalone code completion and generation, sometimes even surpassing human performance, yet their effectiveness diminishes in repository-level settings where cross-file dependencies and structural context are essential. Existing Retrieval-Augmented Generation (RAG) approaches often borrow strategies from NLP, relying on chunking-based indexing and similarity-based retrieval. Chunking results in the loss of coherence between code units and overlooks structural relationships, while similarity-driven methods frequently miss functionally relevant dependencies such as helper functions, classes, or global variables. To address these limitations, we present Hydra, a repository-level code generation framework that treats code as structured code rather than natural language. Our approach introduces (i) a structure-aware indexing strategy that represents repositories as hierarchical trees of functions, classes, and variables, preserving code structure and dependencies, (ii) a lightweight dependency-aware retriever (DAR) that explicitly identifies and retrieves the true dependencies required by a target function, and (iii) a hybrid retrieval mechanism that combines DAR with similarity-based retrieval to provide both essential building blocks and practical usage examples. Extensive experiments on the challenging DevEval and RepoExec benchmarks, both requiring function implementation from real-world repositories with complex large repository context, show that Hydra achieves state-of-the-art performance across open- and closed-source CodeLLMs. Notably, our method establishes a new state of the art in repository-level code generation, surpassing strongest baseline by over 5% in Pass@1 and even enabling smaller models to match or exceed the performance of much larger ones that rely on existing retrievers.<\/jats:p>","DOI":"10.1145\/3797144","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"366-389","source":"Crossref","is-referenced-by-count":0,"title":["Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0000-2859-7038","authenticated-orcid":false,"given":"Minh","family":"Le-Anh","sequence":"first","affiliation":[{"name":"FPT Software AI Center, Hanoi, Vietnam"},{"name":"Hanoi University of Science and Technology, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-4724-9050","authenticated-orcid":false,"given":"Huyen","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0001-4175-7347","authenticated-orcid":false,"given":"An Khanh","family":"Tran","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0005-8895-6051","authenticated-orcid":false,"given":"Nam","family":"Le Hai","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-0011-5137","authenticated-orcid":false,"given":"Linh Ngo","family":"Van","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-1984-4329","authenticated-orcid":false,"given":"Nghi D.Q.","family":"Bui","sequence":"additional","affiliation":[{"name":"FPT Software AI Center, Hanoi, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-5044-1582","authenticated-orcid":false,"given":"Bach","family":"Le","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_1_2_1","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie J. Cai Michael Terry Quoc Le and Charles Sutton. 2021. Program Synthesis with Large Language Models. arXiv:2108.07732 [cs.PL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_2_1_3_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id= oXYZJXDdo7","author":"Cao Bowen","year":"2024","unstructured":"Bowen Cao, Deng Cai, Leyang Cui, Xuxin Cheng, Wei Bi, Yuexian Zou, and Shuming Shi. 2024. Retrieval is Accurate Generation. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id= oXYZJXDdo7"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2023.3267446"},{"key":"e_1_2_1_5_1","volume-title":"CodeT: Code Generation with Generated Tests. In The Eleventh International Conference on Learning Representations. https: \/\/openreview.net\/forum?id=ktrw68Cmu9c","author":"Chen Bei","year":"2023","unstructured":"Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023. CodeT: Code Generation with Generated Tests. In The Eleventh International Conference on Learning Representations. https: \/\/openreview.net\/forum?id=ktrw68Cmu9c"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.137"},{"key":"e_1_2_1_7_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 [cs.LG] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-020-09863-2"},{"key":"e_1_2_1_9_1","unstructured":"Ken Deng Jiaheng Liu He Zhu Congnan Liu Jingxin Li Jiakai Wang Peng Zhao Chenchen Zhang Yanan Wu Xueqiao Yin Yuanxing Zhang Wenbo Su Bangyu Xiang Tiezheng Ge and Bo Zheng. 2024. R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models. arXiv:2406.01359 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2406.01359"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","first-page":"3433","DOI":"10.63317\/55z6pyuui3ij","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024","author":"Ding Yangruibo","year":"2024","unstructured":"Yangruibo Ding, Zijian Wang, Wasi Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang. 2024. CoCoMIC: Code Completion by Jointly Modeling In-file and Cross-file Context. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue (Eds.). ELRA and ICCL, Torino, Italia, 3433-3445. https:\/\/2.zoppoz.workers.dev:443\/https\/aclanthology.org\/2024.lrec-main.305\/"},{"key":"e_1_2_1_12_1","volume-title":"Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang.","author":"Ding Yangruibo","year":"2023","unstructured":"Yangruibo Ding, Zijian Wang, Wasi Uddin Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang. 2023. CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=wgDcbBMSfh"},{"key":"e_1_2_1_13_1","volume-title":"Robert Osazuwa Ness, and Jonathan Larson","author":"Edge Darren","year":"2025","unstructured":"Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2025. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2404.16130"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_2_1_15_1","unstructured":"Yunfan Gao Yun Xiong Xinyu Gao Kangxiang Jia Jinliu Pan Yuxi Bi Yi Dai Jiawei Sun Meng Wang and Haofen Wang. 2024. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2312.10997"},{"key":"e_1_2_1_16_1","unstructured":"Wenchao Gu Juntao Chen Yanlin Wang Tianyue Jiang Xingzhe Li Mingwei Liu Xilin Liu Yuchi Ma and Zibin Zheng. 2025. What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond. arXiv:2503.20589 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2503.20589"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_2_1_18_1","volume-title":"GraphCodeBERT: Pre-training Code Representations with Data Flow. In International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=jLoC4ez43PZ","author":"Guo Daya","year":"2021","unstructured":"Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie LIU, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021. GraphCodeBERT: Pre-training Code Representations with Data Flow. In International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=jLoC4ez43PZ"},{"key":"e_1_2_1_19_1","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y. Wu Y. K. Li Fuli Luo Yingfei Xiong and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming -The Rise of Code Intelligence. arXiv:2401.14196 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2401.14196"},{"key":"e_1_2_1_20_1","volume-title":"Davide Di Ruscio, and Rick Kazman","author":"Hai Nam Le","year":"2025","unstructured":"Nam Le Hai, Anh M. T. Bui, Phuong T. Nguyen, Davide Di Ruscio, and Rick Kazman. 2025. Detection of Technical Debt in Java Source Code. arXiv:2411.05457 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2411.05457"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3643787.3648044"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/"},{"key":"e_1_2_1_23_1","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=sD93GOzH3i5"},{"key":"e_1_2_1_24_1","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu et al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186 (2024)."},{"key":"e_1_2_1_25_1","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu Kai Dang Yang Fan Yichang Zhang An Yang Rui Men Fei Huang Bo Zheng Yibo Miao Shanghaoran Quan Yunlong Feng Xingzhang Ren Xuancheng Ren Jingren Zhou and Junyang Lin. 2024. Qwen2.5-Coder Technical Report. arXiv:2409.12186 [cs.CL] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2409.12186"},{"key":"e_1_2_1_26_1","unstructured":"Aaron Hurst Adam Lerer Adam P Goucher Adam Perelman Aditya Ramesh Aidan Clark AJ Ostrow Akila Welihinda Alan Hayes Alec Radford et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)."},{"key":"e_1_2_1_27_1","unstructured":"Hamel Husain Ho-Hsiang Wu Tiferet Gazit Miltiadis Allamanis and Marc Brockschmidt. 2020. CodeSearchNet Challenge: Evaluating the State of Semantic Code Search. arXiv:1909.09436 [cs.LG] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/1909.09436"},{"key":"e_1_2_1_28_1","volume-title":"Forty-first International Conference on Machine Learning. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=kXHgEYFyf3","author":"Jain Naman","year":"2024","unstructured":"Naman Jain, Manish Shetty, Tianjun Zhang, King Han, Koushik Sen, and Ion Stoica. 2024. R2E: Turning any Github Repository into a Programming Agent Environment. In Forty-first International Conference on Machine Learning. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=kXHgEYFyf3"},{"key":"e_1_2_1_29_1","volume-title":"Silvio Savarese, and Steven Hoi.","author":"Le Hung","year":"2022","unstructured":"Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, and Steven Hoi. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. In Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum? id=WaGvb7OzySA"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.214"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.1072"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3671016.3674819"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-023-10297-9"},{"key":"e_1_2_1_34_1","volume-title":"Hongwei Chen, Chengpeng Wang, and Gang Fan.","author":"Liang Ming","year":"2024","unstructured":"Ming Liang, Xiaoheng Xie, Gehao Zhang, Xunjin Zheng, Peng Di, wei jiang, Hongwei Chen, Chengpeng Wang, and Gang Fan. 2024. REPOFUSE: Repository-Level Code Completion with Fused Dual Context. arXiv:2402.14323 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2402.14323"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3486195"},{"key":"e_1_2_1_36_1","volume-title":"RepoBench: Benchmarking Repository-Level Code Auto- Completion Systems. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/ forum?id=pPjZIOuQuF","author":"Liu Tianyang","year":"2024","unstructured":"Tianyang Liu, Canwen Xu, and Julian McAuley. 2024. RepoBench: Benchmarking Repository-Level Code Auto- Completion Systems. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/ forum?id=pPjZIOuQuF"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695054"},{"key":"e_1_2_1_38_1","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allal Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei Tianyang Liu Max Tian Denis Kocetkov Arthur Zucker Younes Belkada Zijian Wang Qian Liu Dmitry Abulkhanov Indraneil Paul Zhuang Li Wen-Ding Li Megan Risdal Jia Li Jian Zhu Terry Yue Zhuo Evgenii Zheltonozhskii Nii Osae Osae Dade Wenhao Yu Lucas Krau\u00df Naman Jain Yixuan Su Xuanli He Manan Dey Edoardo Abati Yekun Chai Niklas Muennighoff Xiangru Tang Muhtasham Oblokulov Christopher Akiki Marc Marone Chenghao Mou Mayank Mishra Alex Gu Binyuan Hui Tri Dao Armel Zebaze Olivier Dehaene Nicolas Patry Canwen Xu Julian McAuley Han Hu Torsten Scholak Sebastien Paquet Jennifer Robinson Carolyn Jane Anderson Nicolas Chapados Mostofa Patwary Nima Tajbakhsh Yacine Jernite Carlos Mu\u00f1oz Ferrandis Lingming Zhang Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2024. StarCoder 2 and The Stack v2: The Next Generation. arXiv:2402.19173 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2402.19173"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.316"},{"key":"e_1_2_1_40_1","volume-title":"Nam Le Hai, Tien-Thong Doan, Nam V. Nguyen, Quang Pham, and Nghi D. Q. Bui.","author":"Nguyen Dung Manh","year":"2025","unstructured":"Dung Manh Nguyen, Thang Chau Phan, Nam Le Hai, Tien-Thong Doan, Nam V. Nguyen, Quang Pham, and Nghi D. Q. Bui. 2025. CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs. In The Thirteenth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id= CahIEKCu5Q"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.naacl-short.12"},{"key":"e_1_2_1_42_1","unstructured":"OpenAI. 2025. Introducing GPT-4.1 in the API. https:\/\/2.zoppoz.workers.dev:443\/https\/openai.com\/index\/gpt-4-1\/. Accessed: 2025-05-20."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_2_1_44_1","volume-title":"Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al.","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPC66645.2025.00017"},{"key":"e_1_2_1_46_1","volume-title":"RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=GN921JHCRw","author":"Sarthi Parth","year":"2024","unstructured":"Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024. RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. In The Twelfth International Conference on Learning Representations. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=GN921JHCRw"},{"key":"e_1_2_1_47_1","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M Dai Anja Hauth Katie Millican et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)."},{"key":"e_1_2_1_48_1","volume-title":"Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al.","author":"Team Gemini","year":"2024","unstructured":"Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530 (2024)."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.emnlp-main.385"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2026.116001"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE55347"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-naacl.176"},{"key":"e_1_2_1_54_1","volume-title":"Repoformer: Selective Retrieval for Repository-Level Code Completion. In Forty-first International Conference on Machine Learning, ICML 2024","author":"Wu Di","year":"2024","unstructured":"Di Wu, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan, and Xiaofei Ma. 2024. Repoformer: Selective Retrieval for Repository-Level Code Completion. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=moyG54Okrj"},{"key":"e_1_2_1_55_1","unstructured":"Yiqing Xie Alex Xie Divyanshu Sheth Pengfei Liu Daniel Fried and Carolyn Rose. 2024. CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks. arXiv:2404.00566 [cs.SE] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2404.00566"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623316"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_2_1_58_1","unstructured":"Penghao Zhao Hailin Zhang Qinhan Yu Zhengren Wang Yunteng Geng Fangcheng Fu Ling Yang Wentao Zhang Jie Jiang and Bin Cui. 2024. Retrieval-Augmented Generation for AI-Generated Content: A Survey. arXiv:2402.19473 [cs.CV] https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2402.19473"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797144","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:59:45Z","timestamp":1782845985000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3797144"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":58,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797144"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3797144","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}