{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:57:36Z","timestamp":1782845856222,"version":"3.54.5"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"name":"National Key Research and Development Program of China","award":["No. 2024YFB4506300"],"award-info":[{"award-number":["No. 2024YFB4506300"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["Grant Nos. 62322208 and 12411530122"],"award-info":[{"award-number":["Grant Nos. 62322208 and 12411530122"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>False-positive bug reports represent a significant yet underexplored challenge in the development and maintenance of the Linux kernel.  \nThey occur when correct system behavior is mistakenly flagged as a defect, consuming developer effort without leading to actual code improvements. Such reports can mislead developers, waste debugging resources, and delay the resolution of real bugs.  \nIn this paper, we present the first comprehensive empirical study of false-positive bug reports in the Linux kernel.  \nWe manually construct a dataset of 2,006 bug reports comprising 1,509 genuine bugs and 497 false positives collected from Bugzilla and Syzkaller.  \nOur analysis indicates that false positives demand effort comparable to real bugs, often requiring extended discussions and non-trivial closure time.  \nThey occur in several components, especially File Systems and Drivers, mainly due to external dependencies and semantic misunderstandings.  \nTo address this challenge, we evaluate large language models (LLMs) for automated false-positive bug report mitigation.  \nAmong various prompting strategies, retrieval-augmented generation (RAG) performs best, achieving 91% recall and an F1 score of 88%.  \nThese findings highlight the non-negligible cost of false positive bug reports and show the promise of LLMs for more efficient false positive mitigation in the Linux kernel.<\/jats:p>","DOI":"10.1145\/3797082","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"1196-1218","source":"Crossref","is-referenced-by-count":0,"title":["Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0008-4043-3675","authenticated-orcid":false,"given":"Jiashuo","family":"Tian","sequence":"first","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-2004-0902","authenticated-orcid":false,"given":"Dong","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-0759-940X","authenticated-orcid":false,"given":"Chen","family":"Yang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-6953-8369","authenticated-orcid":false,"given":"Haichi","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-6173-8170","authenticated-orcid":false,"given":"Zan","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-3056-9962","authenticated-orcid":false,"given":"Junjie","family":"Chen","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2025. Apache. https:\/\/2.zoppoz.workers.dev:443\/https\/www.apache.org."},{"key":"e_1_2_1_2_1","unstructured":"2025. Bugzilla. https:\/\/2.zoppoz.workers.dev:443\/https\/bugzilla.kernel.org."},{"key":"e_1_2_1_3_1","unstructured":"2025. Deepseek. https:\/\/2.zoppoz.workers.dev:443\/https\/www.deepseek.com."},{"key":"e_1_2_1_4_1","unstructured":"2025. Eclipse. https:\/\/2.zoppoz.workers.dev:443\/https\/www.eclipse.org."},{"key":"e_1_2_1_5_1","unstructured":"2025. Kernel Bugzilla Components. https:\/\/2.zoppoz.workers.dev:443\/https\/bugzilla.kernel.org\/describecomponents.cgi."},{"key":"e_1_2_1_6_1","unstructured":"2025. Mozilla. https:\/\/2.zoppoz.workers.dev:443\/https\/www.mozilla.org."},{"key":"e_1_2_1_7_1","unstructured":"2025. Qwen. https:\/\/2.zoppoz.workers.dev:443\/https\/chat.qwen.ai."},{"key":"e_1_2_1_8_1","unstructured":"2025. Syzkaller. https:\/\/2.zoppoz.workers.dev:443\/https\/syzkaller.appspot.com\/upstream."},{"key":"e_1_2_1_9_1","unstructured":"2025. Trinity: Linux system call fuzzer. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/kernelslacker\/trinity.."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642937.2642990"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1117696.1117704"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3643787.3648043"},{"key":"e_1_2_1_13_1","first-page":"209","volume-title":"OSDI","volume":"8","author":"Cadar Cristian","year":"2008","unstructured":"Cristian Cadar, Daniel Dunbar, Dawson R Engler, et al. 2008. Klee: unassisted and automatic generation of high-coverage tests for complex systems programs.. In OSDI, Vol. 8. 209-224."},{"key":"e_1_2_1_14_1","first-page":"2359","volume-title":"Proceedings of the ACM on Software Engineering 2, FSE","author":"Chen Junjie","year":"2025","unstructured":"Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, and Jun Sun. 2025. De-duplicating Silent Compiler Bugs via Deep Semantic Representation. Proceedings of the ACM on Software Engineering 2, FSE (2025), 2359-2381."},{"key":"e_1_2_1_15_1","volume-title":"Dominance statistics: Ordinal analyses to answer ordinal questions. Psychological bulletin 114, 3","author":"Cliff Norman","year":"1993","unstructured":"Norman Cliff. 1993. Dominance statistics: Ordinal analyses to answer ordinal questions. Psychological bulletin 114, 3 (1993), 494."},{"key":"e_1_2_1_16_1","volume-title":"A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1","author":"Cohen Jacob","year":"1960","unstructured":"Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement 20, 1 (1960), 37-46."},{"key":"e_1_2_1_17_1","volume-title":"Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative sociology 13, 1","author":"Corbin Juliet M","year":"1990","unstructured":"Juliet M Corbin and Anselm Strauss. 1990. Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative sociology 13, 1 (1990), 3-21."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE52982.2021.00030"},{"key":"e_1_2_1_19_1","volume-title":"Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1","author":"Gao Yunfan","year":"2023","unstructured":"Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1 (2023)."},{"key":"e_1_2_1_20_1","unstructured":"Jiawei Gu Xuhui Jiang Zhichao Shi Hexiang Tan Xuehao Zhai Chengjin Xu Wei Li Yinghan Shen Shengjie Ma Honghao Liu et al. 2024. A survey on llm-as-a-judge. arXiv preprint arXiv:2411.15594 (2024)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE5003.2020.00026"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2013.6606585"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485832.3488011"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372297.3417240"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00194"},{"key":"e_1_2_1_26_1","volume-title":"Yeongjin Jang, Insik Shin, and Byoungyoung Lee.","author":"Kim Kyungtae","year":"2020","unstructured":"Kyungtae Kim, Dae R Jeong, Chung Hwan Kim, Yeongjin Jang, Insik Shin, and Byoungyoung Lee. 2020. HFL: Hybrid Fuzzing on the Linux Kernel.. In NDSS."},{"key":"e_1_2_1_27_1","volume-title":"Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa.","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems 35 (2022), 22199-22213."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-024-10459-3"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-024-10502-3"},{"key":"e_1_2_1_30_1","volume-title":"International Conference on Product-Focused Software Process Improvement. Springer, 497-507","author":"Laiq Muhammad","year":"2022","unstructured":"Muhammad Laiq, Nauman bin Ali, J\u00fcrgen B\u00f6stler, and Emelie Engstr\u00f6m. 2022. Early identification of invalid bug reports in industrial settings-a case study. In International Conference on Product-Focused Software Process Improvement. Springer, 497-507."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2023.107305"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134072"},{"key":"e_1_2_1_33_1","unstructured":"Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)."},{"key":"e_1_2_1_34_1","volume-title":"Interrater reliability: the kappa statistic. Biochemia medica 22, 3","author":"McHugh Mary L","year":"2012","unstructured":"Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276-282."},{"key":"e_1_2_1_35_1","volume-title":"Mann-whitney U test. The Corsini encyclopedia of psychology","author":"McKnight Patrick E","year":"2010","unstructured":"Patrick E McKnight and Julius Najab. 2010. Mann-whitney U test. The Corsini encyclopedia of psychology (2010), 1-1."},{"key":"e_1_2_1_36_1","first-page":"919","volume-title":"27th USENIX Security Symposium (USENIX Security 18)","author":"Mu Dongliang","year":"2018","unstructured":"Dongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu, Xinyu Xing, Bing Mao, and Gang Wang. 2018. Understanding the reproducibility of crowd-reported security vulnerabilities. In 27th USENIX Security Symposium (USENIX Security 18). 919-936."},{"key":"e_1_2_1_37_1","volume-title":"Network and Distributed Systems Security Symposium (NDSS).","author":"Mu Dongliang","year":"2022","unstructured":"Dongliang Mu, Yuhang Wu, Yueqi Chen, Zhenpeng Lin, Chensheng Yu, Xinyu Xing, and Gang Wang. 2022. An in-depth analysis of duplicated linux kernel bug reports. In Network and Distributed Systems Security Symposium (NDSS)."},{"key":"e_1_2_1_38_1","first-page":"729","volume-title":"27th USENIX Security Symposium (USENIX Security 18)","author":"Pailoor Shankara","year":"2018","unstructured":"Shankara Pailoor, Andrew Aday, and Suman Jana. 2018. {MoonShine}: Optimizing {OS} fuzzer seed selection with trace distillation. In 27th USENIX Security Symposium (USENIX Security 18). 729-743."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11334-017-0294-1"},{"key":"e_1_2_1_40_1","first-page":"2559","volume-title":"29th USENIX Security Symposium (USENIX Security 20)","author":"Peng Hui","year":"2020","unstructured":"Hui Peng and Mathias Payer. 2020. {USBFuzz}: A framework for fuzzing {USB} drivers by device emulation. In 29th USENIX Security Symposium (USENIX Security 20). 2559-2575."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.2307\/1402731"},{"key":"e_1_2_1_42_1","unstructured":"Jeanine Romano Jeffrey D Kromrey Jesse Coraggio and Jeff Skowronek. 2006. Appropriate statistics for ordinal level data: Should we really be using t-test and Cohen'sd for evaluating group differences on the NSSE and other surveys. In annual meeting of the Florida Association of Institutional Research Vol. 177."},{"key":"e_1_2_1_43_1","volume-title":"26th USENIX security symposium (USENIX Security 17). 167-182.","author":"Schumilo Sergej","unstructured":"Sergej Schumilo, Cornelius Aschermann, Robert Gawlik, Sebastian Schinzel, and Thorsten Holz. 2017. {kAFL}:{Hardware-Assisted} feedback fuzzing for {OS} kernels. In 26th USENIX security symposium (USENIX Security 17). 167-182."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468591"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2019.23176"},{"key":"e_1_2_1_46_1","unstructured":"Donna Spencer. 2009. Card sorting: Designing usable categories. Rosenfeld Media."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477132.3483547"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST.2011.43"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 1630-1644","author":"Tan Xin","year":"2023","unstructured":"Xin Tan, Yuan Zhang, Jiadong Lu, Xin Xiong, Zhuang Liu, and Min Yang. 2023. Syzdirect: Directed greybox fuzzing for linux kernel. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 1630-1644."},{"key":"e_1_2_1_50_1","volume-title":"2017 IEEE international conference on software maintenance and evolution (ICSME). IEEE, 534-538","author":"Terdchanakul Pannavat","year":"2017","unstructured":"Pannavat Terdchanakul, Hideaki Hata, Passakorn Phannachitta, and Kenichi Matsumoto. 2017. Bug or not? bug report classification using n-gram idf. In 2017 IEEE international conference on software maintenance and evolution (ICSME). IEEE, 534-538."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-023-10366-z"},{"key":"e_1_2_1_52_1","first-page":"2741","volume-title":"30th USENIX Security Symposium (USENIX Security 21)","author":"Wang Daimeng","year":"2021","unstructured":"Daimeng Wang, Zheng Zhang, Hang Zhang, Zhiyun Qian, Srikanth V Krishnamurthy, and Nael Abu-Ghazaleh. 2021. {SyzVegas}: Beating kernel fuzzing odds with reinforcement learning. In 30th USENIX Security Symposium (USENIX Security 21). 2741-2758."},{"key":"e_1_2_1_53_1","volume-title":"Denny Zhou, et al.","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824-24837."},{"key":"e_1_2_1_54_1","volume-title":"Wilcoxon signed-rank test","author":"Woolson Robert F","year":"2007","unstructured":"Robert F Woolson. 2007. Wilcoxon signed-rank test. Wiley encyclopedia of clinical trials (2007), 1-3."},{"key":"e_1_2_1_55_1","first-page":"2397","volume-title":"29th USENIX Security Symposium (USENIX Security 20)","author":"Xu Zhengzi","year":"2020","unstructured":"Zhengzi Xu, Yulong Zhang, Longri Zheng, Liangzhao Xia, Chenfu Bao, Zhi Wang, and Yang Liu. 2020. Automatic hot patch generation for android kernels. In 29th USENIX Security Symposium (USENIX Security 20). 2397-2414."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598087"},{"key":"e_1_2_1_57_1","volume-title":"Advancing Code Coverage: Incorporating Program Analysis with Large Language Models. ACM Transactions on Software Engineering and Methodology","author":"Yang Chen","year":"2025","unstructured":"Chen Yang, Junjie Chen, Bin Lin, Ziqi Wang, and Jianyi Zhou. 2025. Advancing Code Coverage: Incorporating Program Analysis with Large Language Models. ACM Transactions on Software Engineering and Methodology (2025)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE63991.2025.00233"},{"key":"e_1_2_1_59_1","volume-title":"WiseUT: An Intelligent Framework for Unit Test Generation. In 2026 IEEE\/ACM 48th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion).","author":"Yang Chen","year":"2026","unstructured":"Chen Yang, Ziqi Wang, Lin Yang, Dong Wang, Shutao Gao, Yanjie Jiang, and Junjie Chen. 2026. WiseUT: An Intelligent Framework for Unit Test Generation. In 2026 IEEE\/ACM 48th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion)."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE63991.2025.00250"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695529"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/WCRE.2012.32"},{"key":"e_1_2_1_63_1","volume-title":"Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv preprint arXiv:2506.05176","author":"Zhang Yanzhao","year":"2025","unstructured":"Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, et al. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv preprint arXiv:2506.05176 (2025)."},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1002\/smr.1770"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797082","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:57:28Z","timestamp":1782842248000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3797082"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":64,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797082"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3797082","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}