{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:01:32Z","timestamp":1782846092097,"version":"3.54.5"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2525507"],"award-info":[{"award-number":["2525507"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Data leakage remains a critical yet under-diagnosed issue in machine learning pipelines, leading to inflated results and unreliable deployments. Existing detection approaches rely on static rules that often miss open-coded manipulations and fail to capture the diversity of real-world notebooks. This paper introduces a novel methodology that integrates static slicing with large language models (LLMs) to improve leakage detection. We use a Datalog-based static analysis that isolates compact, provenance-aware slices corresponding to model training and evaluation pairs, and we pair these with structured LLM prompts that guide step-by-step reasoning about potential leakage for each isolated slice. Evaluated on a curated benchmark of Python notebooks from Kaggle and GitHub, our approach achieves state-of-the-art performance in both preprocessing and overlap leakage detection, improving F1 scores over the previous state-of-the-art by 22% for preprocessing leakage and 15% for overlap leakage. Beyond these improvements, our slicing-based methodology substantially outperforms end-to-end prompting, demonstrating that precise program slicing is key to enabling LLMs to reliably detect leakage. Our findings highlight the effectiveness of combining program slicing and prompt engineering for data leakage detection and establish the first systematic LLM-based solution for detecting data leakage in machine learning code.<\/jats:p>","DOI":"10.1145\/3808199","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"4371-4392","source":"Crossref","is-referenced-by-count":0,"title":["Improving Data Leakage Detection in Machine Learning Notebooks through Static Slicing and Structured LLM Prompts"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0004-8245-9022","authenticated-orcid":false,"given":"Taha","family":"Draoui","sequence":"first","affiliation":[{"name":"University of Michigan-Flint, Flint, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-6010-7561","authenticated-orcid":false,"given":"Mohamed Wiem","family":"Mkaouer","sequence":"additional","affiliation":[{"name":"University of Michigan-Flint, Flint, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-8838-4074","authenticated-orcid":false,"given":"Christian","family":"Newman","sequence":"additional","affiliation":[{"name":"Rochester Institute of Technology, Rochester, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Foundations of Databases: The Logical Level","author":"Abiteboul Serge","unstructured":"Serge Abiteboul, Richard Hull, and Victor Vianu. 1995. Foundations of Databases: The Logical Level (1st ed.). Addison- Wesley Longman Publishing Co., Inc., USA.","edition":"1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593434.3593468"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2771284.2771287"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/SANER64311.2025.00089"},{"key":"e_1_2_1_5_1","volume-title":"2025 IEEE International Conference on Collaborative Advances in Software and COmputiNg (CASCON). IEEE, 389-398","author":"AlOmar Eman Abdullah","year":"2025","unstructured":"Eman Abdullah AlOmar, Luo Xu, Sofia Martinez, Anthony Peruma, Mohamed Wiem Mkaouer, Christian D Newman, and Ali Ouni. 2025. ChatGPT for code refactoring: Analyzing topics, interaction, and effective prompts. In 2025 IEEE International Conference on Collaborative Advances in Software and COmputiNg (CASCON). IEEE, 389-398."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.2730"},{"key":"e_1_2_1_7_1","volume-title":"Assessing Large Language Models Effectiveness in Outdated Method Renaming. In International Conference on Service-Oriented Computing. Springer, 253-260","author":"Mrad Ali Ben","year":"2024","unstructured":"Ali Ben Mrad, Abdoul Majid O. Thiombiano, Mohamed Wiem Mkaouer, and Brahim Hnich. 2024. Assessing Large Language Models Effectiveness in Outdated Method Renaming. In International Conference on Service-Oriented Computing. Springer, 253-260."},{"key":"e_1_2_1_8_1","first-page":"1","article-title":"Deepchecks: A library for testing and validating machine learning models and data","volume":"23","author":"Chorev Shir","year":"2022","unstructured":"Shir Chorev, Philip Tannor, Dan Ben Israel, Noam Bressler, Itay Gabbay, Nir Hutnik, Jonatan Liberman, Matan Perlmutter, Yurii Romanyshyn, and Lior Rokach. 2022. Deepchecks: A library for testing and validating machine learning models and data. Journal of Machine Learning Research 23, 285 (2022), 1-6.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/115372.115320"},{"key":"e_1_2_1_10_1","volume-title":"Self-collaboration Code Generation via ChatGPT. arXiv preprint arXiv:2304.07590","author":"Dong Yihong","year":"2023","unstructured":"Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2023. Self-collaboration Code Generation via ChatGPT. arXiv preprint arXiv:2304.07590 (2023)."},{"key":"e_1_2_1_11_1","volume-title":"Abstract interpretation-based data leakage static analysis. arXiv preprint arXiv:2211.16073","author":"Drobnjakovi\u0107 Filip","year":"2022","unstructured":"Filip Drobnjakovi\u0107, Pavle Suboti\u0107, and Caterina Urban. 2022. Abstract interpretation-based data leakage static analysis. arXiv preprint arXiv:2211.16073 (2022)."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 47th IEEE Computer Software and Applications Conference. 1-10","author":"Feng Yunhe","year":"2023","unstructured":"Yunhe Feng, Sreecharan Vanam, Manasa Cherukupally, Weijian Zheng, Meikang Qiu, and Haihua Chen. 2023. Investigating Code Generation Performance of Chat-GPT with Crowdsourcing Social Data. In Proceedings of the 47th IEEE Computer Software and Applications Conference. 1-10."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.4108\/airo.v2i1.3276"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300500"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","first-page":"2438610","DOI":"10.1080\/12460125.2024.2438610","article-title":"From apologies to insights: extracting topics from chatgpt apologetic responses","volume":"34","author":"Hnich Brahim","year":"2025","unstructured":"Brahim Hnich, Ali Ben Mrad, Abdoul Majid O Thiombiano, and Mohamed Wiem Mkaouer. 2025. From apologies to insights: extracting topics from chatgpt apologetic responses. Journal of Decision Systems 34, 1 (2025), 2438610.","journal-title":"Journal of Decision Systems"},{"key":"e_1_2_1_16_1","volume-title":"Improving Data Scientist Efficiency with Provenance. In 2020 IEEE\/ACM 42nd International Conference on Software Engineering (ICSE). 1086-1097","author":"Hu Jingmei","unstructured":"Jingmei Hu, Jiwon Joung, Maia Jacobs, Krzysztof Z. Gajos, and Margo I. Seltzer. 2020. Improving Data Scientist Efficiency with Provenance. In 2020 IEEE\/ACM 42nd International Conference on Software Engineering (ICSE). 1086-1097."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patter.2023.100804"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Enkelejda Kasneci Kathrin Se\u00dfler Stefan K\u00fcchemann Maria Bannert Daryna Dementieva Frank Fischer Urs Gasser Georg Groh Stephan G\u00fcnnemann Eyke H\u00fcllermeier et al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and individual differences 103 (2023) 102274.","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2382577.2382579"},{"key":"e_1_2_1_20_1","volume-title":"Software Refactoring Research with Large Language Models: A Systematic Literature Review. Journal of Systems and Software","author":"Martinez Sofia","year":"2025","unstructured":"Sofia Martinez, Luo Xu, Mariam Elnaggar, and Eman Abdullah Alomar. 2025. Software Refactoring Research with Large Language Models: A Systematic Literature Review. Journal of Systems and Software (2025), 112762."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 33rd Annual International Conference on Computer Science and Software Engineering. 24-33","author":"Nathalia Nascimento","year":"2023","unstructured":"Nascimento Nathalia, Alencar Paulo, and Cowan Donald. 2023. Artificial Intelligence vs. Software Engineers: An Empirical Study on Performance and Efficiency using ChatGPT. In Proceedings of the 33rd Annual International Conference on Computer Science and Software Engineering. 24-33."},{"key":"e_1_2_1_22_1","volume-title":"On the Performance of Large Language Models for Code Change Intent Classification. In IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER).","author":"Oukay Issam","year":"2025","unstructured":"Issam Oukay, Moataz Chouchen, Ali Ouni, and Fatemeh Hendijani Fard. 2025. On the Performance of Large Language Models for Code Change Intent Classification. In IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)."},{"key":"e_1_2_1_23_1","volume-title":"Evaluating and Explaining Large Language Models for Code Using Syntactic Structures. arXiv preprint arXiv:2308.03873","author":"Palacio David N","year":"2023","unstructured":"David N Palacio, Alejandro Velasco, Daniel Rodriguez-Cardenas, Kevin Moran, and Denys Poshyvanyk. 2023. Evaluating and Explaining Large Language Models for Code Using Syntactic Structures. arXiv preprint arXiv:2308.03873 (2023)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1038\/d41586-018-07196-1"},{"key":"e_1_2_1_25_1","volume-title":"From Copilot to Pilot: Towards AI Supported Software Development. arXiv preprint arXiv:2303.04142","author":"Pudari Rohith","year":"2023","unstructured":"Rohith Pudari and Neil A Ernst. 2023. From Copilot to Pilot: Towards AI Supported Software Development. arXiv preprint arXiv:2303.04142 (2023)."},{"key":"e_1_2_1_26_1","volume-title":"Can large language models identify and refactor code clones? An empirical study. Journal of Systems and Software 234","author":"Qian Xing","year":"2026","unstructured":"Xing Qian and Eman Abdullah AlOmar. 2026. Can large language models identify and refactor code clones? An empirical study. Journal of Systems and Software 234 (2026)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/msr52588.2021.00072"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3696630.3728518"},{"key":"e_1_2_1_29_1","volume-title":"Unveiling the potential of large language models in generating semantic and cross-language clones. arXiv preprint arXiv:2309.06424","author":"Roy Palash R","year":"2023","unstructured":"Palash R Roy, Ajmain I Alam, Farouq Al-omari, Banani Roy, Chanchal K Roy, and Kevin A Schneider. 2023. Unveiling the potential of large language models in generating semantic and cross-language clones. arXiv preprint arXiv:2309.06424 (2023)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3173606"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/3565838.3565855"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2594291.2594320"},{"key":"e_1_2_1_33_1","volume-title":"An analysis of the automatic bug fixing performance of chatgpt. arXiv preprint arXiv:2301.08653","author":"Sobania Dominik","year":"2023","unstructured":"Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke. 2023. An analysis of the automatic bug fixing performance of chatgpt. arXiv preprint arXiv:2301.08653 (2023)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","unstructured":"Johannes St\u00fcmpfle Devansh Atray Nasser Jazdi and Michael Weyrich. 2025. Large Language Model assisted Transformation of Software Variants into a Software Product Line. 12-20. doi:10.1109\/ICSR66718.2025.00008 10.1109\/ICSR66718.2025.00008","DOI":"10.1109\/ICSR66718.2025.00008"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510457.3513032"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3696630.3730565"},{"key":"e_1_2_1_37_1","volume-title":"She Elicits Requirements and He Tests: Software Engineering Gender Bias in Large Language Models. arXiv preprint arXiv:2303.10131","author":"Treude Christoph","year":"2023","unstructured":"Christoph Treude and Hideaki Hata. 2023. She Elicits Requirements and He Tests: Software Engineering Gender Bias in Large Language Models. arXiv preprint arXiv:2303.10131 (2023)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/996893.996859"},{"key":"e_1_2_1_39_1","volume-title":"Jesse Spencer- Smith, and Douglas C Schmidt","author":"White Jules","year":"2023","unstructured":"Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer- Smith, and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382 (2023)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2014.44"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556918"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3808199","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3808199","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:04:49Z","timestamp":1782842689000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3808199"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":41,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808199"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3808199","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}