{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:44:40Z","timestamp":1782848680745,"version":"3.54.5"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T00:00:00Z","timestamp":1782777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Jailbreak attacks have been regarded as a crucial threat to LLM-powered software systems. Recent studies indicate the existence of a steering vector within models' internal activations, which can adjust a model's propensity to reject user requests, and thus is regarded as an effective approach for training-free defense. However, attackers may wrap their malicious intentions within a seemingly benign context, which shifts the distribution of harmful prompts toward benign inputs along the steering vector, effectively bypassing existing defense approaches. In this work, we propose a defense framework InDe-LLM based on intention disentangling. By projecting the embedding of inputs into a benign-invariant subspace, we could disentangle the harmful intentions of jailbreak prompts without affecting benign inputs. Next, such disentangled harmful intentions can be easily identified based on LLMs' well-aligned concept of harmfulness, and rejected through activation steering. Our experiments show that InDe-LLM achieves high defense effectiveness, outperforming baselines by 27.2%\u201343.5% across three models and ten attacks while preserving high utility on benign inputs. Moreover, our evaluation demonstrates that it exhibits high transferability to unseen attacks.<\/jats:p>","DOI":"10.1145\/3797136","type":"journal-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:06:14Z","timestamp":1782839174000},"page":"899-919","source":"Crossref","is-referenced-by-count":0,"title":["InDe-LLM: Defending against Jailbreak Attacks in LLM-Powered Systems via Intention Disentangling"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0008-3892-9145","authenticated-orcid":false,"given":"Yujue","family":"Wang","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-7778-4243","authenticated-orcid":false,"given":"Quan","family":"Zhang","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-6446-247X","authenticated-orcid":false,"given":"Chijin","family":"Zhou","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0001-0461-9674","authenticated-orcid":false,"given":"Gwihwan","family":"Go","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-5379-2100","authenticated-orcid":false,"given":"Dalong","family":"Shi","sequence":"additional","affiliation":[{"name":"AVIC International Digital Network Technology Co., Ltd., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-0955-503X","authenticated-orcid":false,"given":"Yu","family":"Jiang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"n. d.]. AIM Prompt Jailbreak Attack. https:\/\/2.zoppoz.workers.dev:443\/https\/oxtia.com\/chatgpt-jailbreak-prompts\/aim-prompt\/."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.14132"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.02151"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.09932"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4322"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.07875"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.14348"},{"key":"e_1_2_1_8_1","volume-title":"Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023","author":"Carlini Nicholas","year":"2023","unstructured":"Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tram\u00e8r, and Ludwig Schmidt. 2023. Are aligned neural networks adversarially aligned?. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 -16, 2023, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.). https:\/\/2.zoppoz.workers.dev:443\/http\/papers.nips.cc\/paper_files\/paper\/2023\/hash\/c1f0b856a35986348ab3414177266f75- Abstract-Conference.html"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2024.FINDINGS-ACL.304"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/SATML64287.2025.00010"},{"key":"e_1_2_1_11_1","volume-title":"Deep reinforcement learning from human preferences. Advances in neural information processing systems 30","author":"Christiano Paul F","year":"2017","unstructured":"Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_1_12_1","volume-title":"Forty-first International Conference on Machine Learning, ICML 2024","author":"Dao Tri","year":"2024","unstructured":"Tri Dao and Albert Gu. 2024. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=ztn8FCR1td"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.08268"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et al. 2024. The llama 3 herd of models. arXiv e-prints (2024) arXiv-2407. doi:10.48550\/arXiv.2407.21783 10.48550\/arXiv.2407.21783","DOI":"10.48550\/arXiv.2407.21783"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2404.04475"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.02355"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2312.00752"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","unstructured":"Melody Y Guan Manas Joglekar Eric Wallace Saachi Jain Boaz Barak Alec Helyar Rachel Dias Andrea Vallone Hongyu Ren Jason Wei et al. 2024. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339 (2024). doi:10.48550\/arXiv.2412.16339 10.48550\/arXiv.2412.16339","DOI":"10.48550\/arXiv.2412.16339"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.03874"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","unstructured":"Binyuan Hui Jian Yang Zeyu Cui Jiaxi Yang Dayiheng Liu Lei Zhang Tianyu Liu Jiajun Zhang Bowen Yu Keming Lu et al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186 (2024). doi:10.48550\/arXiv.2409.12186 10.48550\/arXiv.2409.12186","DOI":"10.48550\/arXiv.2409.12186"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2502.02716"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.00614"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3584700"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1072"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.14477"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.15180"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2409.05907"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.16914"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies. 746-751","author":"Mikolov Tom\u00e1\u0161","year":"2013","unstructured":"Tom\u00e1\u0161 Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies. 746-751."},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022) 27730-27744.","DOI":"10.52202\/068431-2011"},{"key":"e_1_2_1_32_1","volume-title":"Steering llama 2 via contrastive activation addition. arXiv preprint arXiv:2312.06681","author":"Panickssery Nina","year":"2023","unstructured":"Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023. Steering llama 2 via contrastive activation addition. arXiv preprint arXiv:2312.06681 (2023)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.03658"},{"key":"e_1_2_1_34_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2310.03684"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3560260"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474381"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.11755"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.02298"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658644.3670388"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2506.07022"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2408.00118"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.10248"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.14968"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlpmain.585"},{"key":"e_1_2_1_46_1","volume-title":"The Thirteenth International Conference on Learning Representations, ICLR 2025","author":"Wang Xinpeng","year":"2025","unstructured":"Xinpeng Wang, Chengzhi Hu, Paul R\u00f6ttger, and Barbara Plank. 2025. Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.02483"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2310.02949"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2310.02446"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2504.09466"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2310.15140"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2307.15043"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.14857"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3797136","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T18:52:26Z","timestamp":1782845546000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3797136"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,30]]},"references-count":55,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3797136"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3797136","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,30]]}}}