{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T20:10:06Z","timestamp":1756411806150,"version":"3.44.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"DOI":"10.13039\/100007847","name":"Natural Science Foundation of Jilin Province","doi-asserted-by":"crossref","award":["20220101108JC"],"award-info":[{"award-number":["20220101108JC"]}],"id":[{"id":"10.13039\/100007847","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Project of China","award":["2023YFF0714100"],"award-info":[{"award-number":["2023YFF0714100"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>Due to the limitations of imaging sensors, obtaining a medical image that simultaneously captures both functional metabolic data and structural tissue details remains a significant challenge in clinical diagnosis. To address this, Multimodal Medical Image Fusion (MMIF) has emerged as an effective technique for integrating complementary information from multimodal source images, such as CT, PET, and SPECT, which is critical for providing a comprehensive understanding of both anatomical and functional aspects of the human body. One of the key challenges in MMIF is how to exchange and aggregate this multimodal information. This article rethinks MMIF by addressing the harmony of modality gaps and proposes a novel Modality-Aware Interaction Network (MAINet), which leverages cross-modal feature interaction and progressively fuses multiple features in graph space. Specifically, we introduce two key modules: the Cascade Modality Interaction (CMI) module and the Dual-Graph Learning (DGL) module. The CMI module, integrated within a multi-scale encoder with triple branches, facilitates complementary multimodal feature learning and provides beneficial feedback to enhance discriminative feature learning across modalities. In the decoding process, the DGL module aggregates hierarchical features in two distinct graph spaces, enabling global feature interactions. Moreover, the DGL module incorporates a bottom-up guidance mechanism, where deeper semantic features guide the learning of shallower detail features, thus improving the fusion process by enhancing both scale diversity and modality awareness for visual fidelity results. Experimental results on medical image datasets demonstrate the superiority of the proposed method over existing fusion approaches in both subjective and objective evaluations. We also validated the performance of the proposed method in applications such as infrared-visible image fusion and medical image segmentation.<\/jats:p>","DOI":"10.1145\/3731247","type":"journal-article","created":{"date-parts":[[2025,4,18]],"date-time":"2025-04-18T12:09:51Z","timestamp":1744978191000},"page":"1-23","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MAINet: Modality-Aware Interaction Network for Medical Image Fusion"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-5161-3412","authenticated-orcid":false,"given":"Lisi","family":"Wei","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China and College of Artificial Intelligence and Big Data, Hulunbuir University, Hulunbuir, China and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, Changchun, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0001-3979-8801","authenticated-orcid":false,"given":"Libo","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, Changchun, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-8412-4956","authenticated-orcid":false,"given":"Xiaoli","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Jilin University, Changchun, China and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, Changchun, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Joan Bruna Wojciech Zaremba Arthur Szlam and Yann LeCun. 2013. Spectral networks and locally connected networks on graphs. arXiv:1312.6203. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/1312.6203"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2007.12.002"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2022.11.010"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2009.05.003"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.5555\/1248547.1248548"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649466"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2024.106478"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2011.08.002"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3185887"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3349209"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00729"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19797-0_31"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.120301"},{"key":"e_1_3_3_16_2","first-page":"3519","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kornblith Simon","year":"2019","unstructured":"Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019. Similarity of neural network representations revisited. In Proceedings of the International Conference on Machine Learning. PMLR, 3519\u20133529."},{"key":"e_1_3_3_17_2","first-page":"438","volume-title":"Proceedings of the 1998 IEEE Signal Processing Society Workshop on Neural Networks for Signal Processing VIII (Cat. No. 98TH8378)","author":"Lai Shang-Hong","year":"1998","unstructured":"Shang-Hong Lai and Ming Fang. 1998. Adaptive medical image visualization based on hierarchical neural networks and intelligent decision fusion. In Proceedings of the 1998 IEEE Signal Processing Society Workshop on Neural Networks for Signal Processing VIII (Cat. No. 98TH8378). IEEE, 438\u2013447."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2887342"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612135"},{"key":"e_1_3_3_20_2","first-page":"1","article-title":"GeSeNet: A general semantic-guided network with couple mask ensemble for medical image fusion","volume":"35","author":"Li Jiawei","year":"2023","unstructured":"Jiawei Li, Jinyuan Liu, Shihua Zhou, Qiang Zhang, and Nikola K. Kasabov. 2023. GeSeNet: A general semantic-guided network with couple mask ensemble for medical image fusion. IEEE Transactions on Neural Networks and Learning Systems 35 (2023), 1\u201314.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3574136"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3056725"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3043125"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.23919\/ICIF.2017.8009769"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2022.105770"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.109"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3129354"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2018.02.004"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2022.105686"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121789"},{"key":"e_1_3_3_31_2","unstructured":"Thao Nguyen Maithra Raghu and Simon Kornblith. 2020. Do wide and deep networks learn the same things? Uncovering how neural network representations vary with width and depth. arXiv:2010.15327. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2010.15327"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558770"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijleo.2018.12.028"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2018.04.002"},{"key":"e_1_3_3_35_2","first-page":"1","article-title":"Multiscale attention fusion graph network for remote sensing building change detection","volume":"62","author":"Shangguan Yu","year":"2024","unstructured":"Yu Shangguan, Jinjiang Li, Zheng Chen, Lu Ren, and Zhen Hua. 2024. Multiscale attention fusion graph network for remote sensing building change detection. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1\u201318.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2005.859378"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3205747"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.102163"},{"key":"e_1_3_3_39_2","first-page":"1","article-title":"Multimodal medical image fusion algorithm in the era of big data","author":"Tan Wei","year":"2020","unstructured":"Wei Tan, Prayag Tiwari, Hari Mohan Pandey, Catarina Moreira, and Amit Kumar Jaiswal. 2020. Multimodal medical image fusion algorithm in the era of big data. Neural Computing and Applications (2020), 1\u201321.","journal-title":"Neural Computing and Applications"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.101870"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-87193-2_4"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-16443-9_3"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP46576.2022.9897280"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638557"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICOSP.2008.4697288"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-372529-5.00017-2"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2024.108771"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3404660"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3012548"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1049\/el:20000267"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3245607"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3223216"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01501-8"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.06.008"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02109"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3243853"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2019.07.011"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3296745"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00572"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3731247","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T19:55:21Z","timestamp":1756410921000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3731247"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":59,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3731247"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3731247","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-06-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}