{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T22:31:36Z","timestamp":1757629896150,"version":"3.44.0"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"9","funder":[{"name":"Natural Science Foundation of Jiangsu Province, China","award":["BK20241900"],"award-info":[{"award-number":["BK20241900"]}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program","doi-asserted-by":"crossref","award":["2022YFC2405600"],"award-info":[{"award-number":["2022YFC2405600"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276139 and U2001211"],"award-info":[{"award-number":["62276139 and U2001211"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Person image generation is widely used in many fields, but it still faces some challenges. Most of current person image generation methods suffer from an intractable problem of handling the spatial deformation caused by the pose change in the generation process, while convolution-based generative model is not good at handling\u00a0the region-unaligned task. Therefore, we propose a novel UV-space transformation network to implement the primary generation of person image in the UV-space. This framework can effectively avoid the spatial deformation problems in the generation process and instead transfer them to the preceding pose estimation stage. Within the framework, we propose the self-reconstruction-assisted UV texture transformation blocks which aim to exploit the self-reconstruction of source texture map to guide and assist the generation of target UV texture map. In addition, after obtaining the target person image from the generated UV texture map, we use the correlations between the source and generated images to further improve the details of the generated person images. Superior experiment results compared with other state-of-the-art methods demonstrate the effectiveness of the proposed method.<\/jats:p>","DOI":"10.1145\/3749375","type":"journal-article","created":{"date-parts":[[2025,7,22]],"date-time":"2025-07-22T22:20:26Z","timestamp":1753222826000},"page":"1-16","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Source Information-Assisted UV-Space Transformation Network for Person Image Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-4909-0361","authenticated-orcid":false,"given":"Guiyu","family":"Xia","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0000-1085-688X","authenticated-orcid":false,"given":"Zhedong","family":"Jin","sequence":"additional","affiliation":[{"name":"Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET), Nanjing University of Information Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0004-6780-2355","authenticated-orcid":false,"given":"Dongdong","family":"Fang","sequence":"additional","affiliation":[{"name":"Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET), Nanjing University of Information Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-0462-3729","authenticated-orcid":false,"given":"Yubao","family":"Sun","sequence":"additional","affiliation":[{"name":"Jiangsu Collaborative Innovation Center of Atmospheric Environment and Equipment Technology (CICAEET), Nanjing University of Information Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"issue":"11","key":"e_1_3_2_2_2","first-page":"356:1","article-title":"Revolutionizing visuals: The role of generative AI in modern image generation","volume":"20","author":"Bansal Gaurang","year":"2024","unstructured":"Gaurang Bansal, Aditya Nawal, Vinay Chamola, and Norbert Herencsar. 2024. Revolutionizing visuals: The role of generative AI in modern image generation. ACM Transactions on Multimedia Computing, Communications and Applications 20, 11 (2024), 356:1\u2013356:22.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_3_2","first-page":"5968","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Bhunia Ankan Kumar","year":"2023","unstructured":"Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. 2023. Person image synthesis via denoising diffusion model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5968\u20135976."},{"issue":"1","key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1109\/TCSVT.2021.3059706","article-title":"Pman: Progressive multi-attention network for human pose transfer","volume":"32","author":"Chen Baoyu","year":"2021","unstructured":"Baoyu Chen, Yi Zhang, Hongchen Tan, Baocai Yin, and Xiuping Liu. 2021. Pman: Progressive multi-attention network for human pose transfer. IEEE Transactions on Circuits and Systems for Video Technology 32, 1 (2021), 302\u2013314.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"doi-asserted-by":"publisher","key":"e_1_3_2_5_2","DOI":"10.1109\/CVPR.2018.00762"},{"key":"e_1_3_2_6_2","article-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"doi-asserted-by":"publisher","key":"e_1_3_2_7_2","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_2_8_2","first-page":"694","volume-title":"Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916)","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916). Springer, 694\u2013711."},{"doi-asserted-by":"publisher","key":"e_1_3_2_9_2","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_3_2_10_2","first-page":"2039","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Nannan","year":"2024","unstructured":"Nannan Li, Qing Liu, Krishna Kumar Singh, Yilin Wang, Jianming Zhang, Bryan A. Plummer, and Zhe Lin. 2024. Unihuman: A unified model for editing human images in the wild. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2039\u20132048."},{"doi-asserted-by":"publisher","key":"e_1_3_2_11_2","DOI":"10.1016\/j.neucom.2022.09.089"},{"doi-asserted-by":"publisher","key":"e_1_3_2_12_2","DOI":"10.1016\/j.knosys.2021.107024"},{"doi-asserted-by":"publisher","key":"e_1_3_2_13_2","DOI":"10.1109\/TIP.2021.3107235"},{"doi-asserted-by":"publisher","key":"e_1_3_2_14_2","DOI":"10.1109\/CVPR.2016.124"},{"key":"e_1_3_2_15_2","first-page":"851","article-title":"SMPL: A skinned multi-person linear model","volume":"2","author":"Loper Matthew","year":"2023","unstructured":"Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. 2023. SMPL: A skinned multi-person linear model. Seminal Graphics Papers: Pushing the Boundaries 2 (2023), 851\u2013866.","journal-title":"Seminal Graphics Papers: Pushing the Boundaries"},{"doi-asserted-by":"publisher","key":"e_1_3_2_16_2","DOI":"10.1016\/j.knosys.2023.110852"},{"key":"e_1_3_2_17_2","first-page":"930","volume-title":"IEEE Transactions on Multimedia","author":"Ma Liyuan","year":"2023","unstructured":"Liyuan Ma, Kejie Huang, Dongxu Wei, Zhao-Yan Ming, and Haibin Shen. 2023. Fda-gan: Flow-based dual attention gan for human pose transfer. IEEE Transactions on Multimedia 25 (2023), 930\u2013941."},{"key":"e_1_3_2_18_2","article-title":"Pose guided person image generation","volume":"30","author":"Ma Liqian","year":"2017","unstructured":"Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, and Luc Van Gool. 2017. Pose guided person image generation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"doi-asserted-by":"publisher","key":"e_1_3_2_19_2","DOI":"10.1109\/CVPR42600.2020.00513"},{"doi-asserted-by":"publisher","key":"e_1_3_2_20_2","DOI":"10.1007\/978-3-030-01219-9_8"},{"doi-asserted-by":"publisher","key":"e_1_3_2_21_2","DOI":"10.1145\/3306346.3323016"},{"issue":"3","key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3610534","article-title":"Deep learning based occluded person re-identification: A survey","volume":"20","author":"Peng Yunjie","year":"2023","unstructured":"Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu, Zhenan Sun, and Zhiqiang He. 2023. Deep learning based occluded person re-identification: A survey. ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1\u201327.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"doi-asserted-by":"publisher","key":"e_1_3_2_23_2","DOI":"10.1109\/CVPR46437.2021.01065"},{"doi-asserted-by":"publisher","key":"e_1_3_2_24_2","DOI":"10.1109\/CVPR52688.2022.01317"},{"doi-asserted-by":"publisher","key":"e_1_3_2_25_2","DOI":"10.1109\/CVPR42600.2020.00771"},{"key":"e_1_3_2_26_2","article-title":"Improved techniques for training gans","volume":"29","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 29.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"unstructured":"Kripasindhu Sarkar Vladislav Golyanik Lingjie Liu and Christian Theobalt. 2021. Style and pose control for image synthesis of humans from a single monocular view. arXiv:2102.11263. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2102.11263","key":"e_1_3_2_27_2"},{"key":"e_1_3_2_28_2","first-page":"596","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV\u00a0\u201920)","author":"Sarkar Kripasindhu","year":"2020","unstructured":"Kripasindhu Sarkar, Dushyant Mehta, Weipeng Xu, Vladislav Golyanik, and Christian Theobalt. 2020. Neural re-rendering of humans from a single image. In Proceedings of the 16th European Conference on Computer Vision (ECCV\u00a0\u201920). Springer, 596\u2013613."},{"unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/1409.1556","key":"e_1_3_2_29_2"},{"key":"e_1_3_2_30_2","first-page":"717","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Tang Hao","year":"2020","unstructured":"Hao Tang, Song Bai, Li Zhang, Philip H. S. Torr, and Nicu Sebe. 2020. Xinggan for person image generation. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 717\u2013734."},{"doi-asserted-by":"publisher","key":"e_1_3_2_31_2","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_32_2","article-title":"Semantic map guided identity transfer GAN for person re-identification","author":"Wu Tian","year":"2024","unstructured":"Tian Wu, Rongbo Zhu, and Shaohua Wan. 2024. Semantic map guided identity transfer GAN for person re-identification. ACM Transactions on Multimedia Computing, Communications and Applications 20, 11, Article 333 (Sep. 2024), 20.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_33_2","first-page":"3270","article-title":"3D information guided motion transfer via sequential image based human model refinement and face-attention GAN","author":"Xia Guiyu","year":"2023","unstructured":"Guiyu Xia, Dong Luo, Zeyuan Zhang, Yubao Sun, and Qingshan Liu. 2023. 3D information guided motion transfer via sequential image based human model refinement and face-attention GAN. IEEE Transactions on Circuits and Systems for Video Technology 33, 7 (2023), 3270\u20133283.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"doi-asserted-by":"publisher","key":"e_1_3_2_34_2","DOI":"10.1145\/3554739"},{"doi-asserted-by":"publisher","key":"e_1_3_2_35_2","DOI":"10.1109\/CVPR46437.2021.00789"},{"doi-asserted-by":"publisher","key":"e_1_3_2_36_2","DOI":"10.1109\/CVPR.2018.00068"},{"doi-asserted-by":"publisher","key":"e_1_3_2_37_2","DOI":"10.1109\/TIP.2020.3031108"},{"key":"e_1_3_2_38_2","first-page":"161","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhou Xinyue","year":"2022","unstructured":"Xinyue Zhou, Mingyu Yin, Xinyuan Chen, Li Sun, Changxin Gao, and Qingli Li. 2022. Cross attention based style distribution for controllable person image synthesis. In Proceedings of the European Conference on Computer Vision. Springer, 161\u2013178."},{"unstructured":"Zijian Zhou Shikun Liu Xiao Han Haozhe Liu Kam Woh Ng Tian Xie Yuren Cong Hang Li Mengmeng Xu Juan-Manuel P\u00e9rez-R\u00faa et al. 2024. Learning flow fields in attention for controllable person image generation. arXiv:2412.08486. Retrieved from https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2412.08486","key":"e_1_3_2_39_2"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3749375","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T16:02:21Z","timestamp":1757520141000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3749375"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":38,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3749375"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3749375","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-09-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}