{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:07:21Z","timestamp":1782846441569,"version":"3.54.5"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Effective incident management in large-scale IT systems relies on troubleshooting guides (TSGs), but their manual execution is slow and error-prone. While recent advances in LLMs offer promise for automating incident management tasks, existing LLM-based solutions lack specialized support for several key challenges, including managing TSG quality issues, interpreting complex control flow, handling data-intensive queries, and exploiting execution parallelism. We first conducted an empirical study on 92 real-world TSGs, and, guided by our findings, we present , a novel end-to-end agentic framework for troubleshooting guide automation. Our approach features a three-stage workflow: the first stage provides a comprehensive guide together with a tool, TSG Mentor, to assist site reliability engineers (SREs) in improving TSG quality; the second stage performs offline preprocessing using LLMs to extract structured execution directed acyclic graphs (DAGs) from unstructured TSGs and to create dedicated Query Preparation Plugins (QPPs); and the third stage executes online using a DAG-guided scheduler-executor framework with a memory system to ensure correct workflow and support parallel execution of independent steps. Our empirical evaluation on a collection of real-world TSGs and incidents demonstrates that \u00a0 achieves a \u223c94% success rate on GPT-4.1, outperforming baselines with less time and token consumption. Furthermore, it achieves a remarkable execution time reduction of 32.9% to 70.4% for parallelizable TSGs. Our code and sample data are publicly available at https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/microsoft\/StepFly.<\/jats:p>","DOI":"10.1145\/3808143","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"3070-3092","source":"Crossref","is-referenced-by-count":0,"title":["StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0000-5488-7813","authenticated-orcid":false,"given":"Jiayi","family":"Mao","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-4579-3799","authenticated-orcid":false,"given":"Liqun","family":"Li","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-1899-8561","authenticated-orcid":false,"given":"Yanjie","family":"Gao","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"},{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0003-0325-1593","authenticated-orcid":false,"given":"Zegang","family":"Peng","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-8595-5388","authenticated-orcid":false,"given":"Shilin","family":"He","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-1304-6839","authenticated-orcid":false,"given":"Chaoyun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-8698-1860","authenticated-orcid":false,"given":"Si","family":"Qin","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-9893-594X","authenticated-orcid":false,"given":"Samia","family":"Khalid","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-2559-2383","authenticated-orcid":false,"given":"Qingwei","family":"Lin","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-2019-213X","authenticated-orcid":false,"given":"Saravan","family":"Rajmohan","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0008-7545-9461","authenticated-orcid":false,"given":"Sitaram","family":"Lanka","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-9230-2799","authenticated-orcid":false,"given":"Dongmei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.3233\/FAIA241032"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/234313.234418"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3627703.3629553"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417055"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_2_1_6_1","volume-title":"Thinking in JAVA","author":"Eckel Bruce","unstructured":"Bruce Eckel. 2003. Thinking in JAVA. Prentice Hall Professional."},{"key":"e_1_2_1_7_1","volume-title":"MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In International Conference on Learning Representations (ICLR).","author":"Hong Sirui","year":"2024","unstructured":"Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and J\u00fcrgen Schmidhuber. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417054"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639081"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3613891"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3381"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.2307\/2529310"},{"key":"e_1_2_1_14_1","unstructured":"LangChain AI. 2024. LangGraph: State Machines for LLM Applications. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/langchain-ai\/langgraph. Accessed: 2025-05-30."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3689051.3689056"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-emnlp.873"},{"key":"e_1_2_1_17_1","volume-title":"Xiang Yue, and Wenhu Chen.","author":"Li Tianle","year":"2025","unstructured":"Tianle Li, Ge Zhang, Quy Duc Do, Xiang Yue, and Wenhu Chen. 2025. Long-context LLMs Struggle with Long In-context Learning. Transactions on Machine Learning Research (2025)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2025.3592032"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3711896.3737427"},{"key":"e_1_2_1_20_1","volume-title":"Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding. In International Conference on Learning Representations (ICLR). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=cPozlf9OaF","author":"Liu Yanming","year":"2025","unstructured":"Yanming Liu, Xinyue Peng, Jiannan Cao, Shi Bo, Yanxin Shen, Tianyu Du, Sheng Cheng, Xun Wang, Jianwei Yin, and Xuhong Zhang. 2025. Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding. In International Conference on Learning Representations (ICLR). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=cPozlf9OaF"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2663171.2663188"},{"key":"e_1_2_1_22_1","unstructured":"Microsoft. 2025. Kusto Query Language Overview. https:\/\/2.zoppoz.workers.dev:443\/https\/learn.microsoft.com\/en-us\/kusto\/query\/. Accessed: 2026-04-09."},{"key":"e_1_2_1_23_1","unstructured":"MongoDB Inc. 2025. MongoDB: The World's Leading Modern Database. https:\/\/2.zoppoz.workers.dev:443\/https\/www.mongodb.com\/. Accessed: 2025-05-30."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3701716.3715225"},{"key":"e_1_2_1_25_1","volume-title":"Types and programming languages","author":"Pierce Benjamin C","unstructured":"Benjamin C Pierce. 2002. Types and programming languages. MIT press."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","unstructured":"Bo Qiao Liqun Li Xu Zhang Shilin He Yu Kang Chaoyun Zhang Fangkai Yang Hang Dong Jue Zhang Lu Wang Minghua Ma Pu Zhao Si Qin Xiaoting Qin Chao Du Yong Xu Qingwei Lin Saravan Rajmohan and Dongmei Zhang. 2024. TaskWeaver: A Code-First Agent Framework. arXiv:2311.17541 [cs.AI] doi:10.48550\/arXiv.2311.17541 10.48550\/arXiv.2311.17541","DOI":"10.48550\/arXiv.2311.17541"},{"key":"e_1_2_1_27_1","unstructured":"Toran Bruce Richards. 2023. AutoGPT. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/Significant-Gravitas\/AutoGPT. Accessed: 2025-05-30."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3663529.3663841"},{"key":"e_1_2_1_29_1","volume-title":"Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems 36 (NeurIPS). 68539-68551.","author":"Schick Timo","year":"2023","unstructured":"Timo Schick, Jane Dwivedi-Yu, Roberto Dess\u00ec, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems 36 (NeurIPS). 68539-68551."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3558958"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/323779.323736"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.02933"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE51524.2021.9678708"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE52982.2021.00017"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research","volume":"50232","author":"Wang Xingyao","year":"2024","unstructured":"Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. 2024. Executable Code Actions Elicit Better LLM Agents. In Proceedings of the 41st International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 235). 50208-50232."},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Xuezhi Wang and Denny Zhou. 2024. Chain-of-Thought Reasoning Without Prompting. In Advances in Neural Information Processing Systems 37 (NeurIPS).","DOI":"10.52202\/079017-2123"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE62328.2024.00056"},{"key":"e_1_2_1_38_1","volume-title":"Chi, Quoc Le, and Denny Zhou","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35 (NeurIPS)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/502055.502057"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the Conference on Language Modeling (COLM). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=uAjxFFing2","author":"Wu Qingyun","year":"2024","unstructured":"Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. In Proceedings of the Conference on Language Modeling (COLM). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=uAjxFFing2"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.08694"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","unstructured":"Shuyuan Xu Zelong Li Kai Mei and Yongfeng Zhang. 2024. CoRE: LLM as Interpreter for Natural Language Programming Pseudo-Code Programming and Flow Programming of AI Agents. doi:10.48550\/arXiv.2405.06907 arXiv:2405.06907. 10.48550\/arXiv.2405.06907","DOI":"10.48550\/arXiv.2405.06907"},{"key":"e_1_2_1_43_1","volume-title":"ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=WE_vluYUL-X","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations (ICLR). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=WE_vluYUL-X"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE65448.2025.00011"},{"key":"e_1_2_1_45_1","unstructured":"Xuchao Zhang Tanish Mittal Chetan Bansal Rujia Wang Minghua Ma Zhixin Ren Hao Huang and Saravan Rajmohan. 2024. FLASH: A Workflow Automation Agent for Diagnosing Recurring Incidents. Technical Report. Microsoft Research. https:\/\/2.zoppoz.workers.dev:443\/https\/www.microsoft.com\/en-us\/research\/publication\/flash-a-workflow-automation-agent-for-diagnosingrecurring-incidents\/"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3808143","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:14:40Z","timestamp":1782843280000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3808143"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":45,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3808143"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3808143","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}