{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T01:42:40Z","timestamp":1773193360119,"version":"3.50.1"},"reference-count":34,"publisher":"Wiley","issue":"3","license":[{"start":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T00:00:00Z","timestamp":1716940800000},"content-version":"vor","delay-in-days":28,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62272018"],"award-info":[{"award-number":["62272018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Animation &amp;amp; Virtual"],"published-print":{"date-parts":[[2024,5]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>Generating immersive virtual reality avatars is a challenging task in VR\/AR applications, which maps physical human body poses to avatars in virtual scenes for an immersive user experience. However, most existing work is time\u2010consuming and limited by datasets, which does not satisfy immersive and real\u2010time requirements of VR systems. In this paper, we aim to generate 3D real\u2010time virtual reality avatars based on a monocular camera to solve these problems. Specifically, we first design a self\u2010attention distillation network (SADNet) for effective human pose estimation, which is guided by a pre\u2010trained teacher. Secondly, we propose a lightweight pose mapping method for human avatars that utilizes the camera model to map 2D poses to 3D avatar keypoints, generating real\u2010time human avatars with pose consistency. Finally, we integrate our framework into a VR system, displaying generated 3D pose\u2010driven avatars on Helmet\u2010Mounted Display devices for an immersive user experience. We evaluate SADNet on two publicly available datasets. Experimental results show that SADNet achieves a state\u2010of\u2010the\u2010art trade\u2010off between speed and accuracy. In addition, we conducted a user experience study on the performance and immersion of virtual reality avatars. Results show that pose\u2010driven 3D human avatars generated by our method are smooth and attractive.<\/jats:p>","DOI":"10.1002\/cav.2233","type":"journal-article","created":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T08:04:13Z","timestamp":1716969853000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["SADNet: Generating immersive virtual reality avatars by real\u2010time monocular pose estimation"],"prefix":"10.1002","volume":"35","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-9127-9219","authenticated-orcid":false,"given":"Ling","family":"Jiang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-7253-4998","authenticated-orcid":false,"given":"Yuan","family":"Xiong","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0007-9056-0835","authenticated-orcid":false,"given":"Qianqian","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-2843-7283","authenticated-orcid":false,"given":"Tong","family":"Chen","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-6572-8471","authenticated-orcid":false,"given":"Wei","family":"Wu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-5825-7517","authenticated-orcid":false,"given":"Zhong","family":"Zhou","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems Beihang University Beijing China"},{"name":"Zhongguancun Laboratory Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2024,5,29]]},"reference":[{"key":"e_1_2_9_2_1","doi-asserted-by":"crossref","unstructured":"TangMT ZhuVL PopescuV.Alterecho: loose avatar\u2010streamer coupling for expressive vtubing. 2021 IEEE international symposium on mixed and augmented reality (ISMAR). IEEE 128\u2010137.2021.","DOI":"10.1109\/ISMAR52148.2021.00027"},{"key":"e_1_2_9_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417768"},{"key":"e_1_2_9_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNSRE.2022.3156884"},{"key":"e_1_2_9_5_1","doi-asserted-by":"crossref","unstructured":"ZhangY LiZ AnL LiM YuT LiuY.Lightweight multi\u2010person total motion capture using sparse multi\u2010view cameras. Proceedings of the IEEE\/CVF International Conference on Computer Vision 5560\u20105569.2021.","DOI":"10.1109\/ICCV48922.2021.00551"},{"key":"e_1_2_9_6_1","doi-asserted-by":"crossref","unstructured":"CudeiroD BolkartT LaidlawC RanjanA BlackMJ.Capture learning and synthesis of 3D speaking styles. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 10101\u201010111.2019.","DOI":"10.1109\/CVPR.2019.01034"},{"key":"e_1_2_9_7_1","doi-asserted-by":"crossref","unstructured":"ShaoR ZhangH ZhangH et al.Doublefield: bridging the neural surface and radiance fields for high\u2010fidelity human reconstruction and rendering. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 15872\u201015882.2022.","DOI":"10.1109\/CVPR52688.2022.01541"},{"key":"e_1_2_9_8_1","doi-asserted-by":"crossref","unstructured":"YuT ZhengZ GuoK LiuP DaiQ LiuY.Function4d: real\u2010time human volumetric capture from very sparse consumer rgbd sensors. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 5746\u20105756.2021.","DOI":"10.1109\/CVPR46437.2021.00569"},{"key":"e_1_2_9_9_1","doi-asserted-by":"crossref","unstructured":"LiZ YuT ZhengZ GuoK LiuY.Posefusion: pose\u2010guided selective fusion for single\u2010view human volumetric capture. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 14162\u201014172.2021.","DOI":"10.1109\/CVPR46437.2021.01394"},{"key":"e_1_2_9_10_1","doi-asserted-by":"crossref","unstructured":"KocabasM AthanasiouN BlackMJ.Vibe: video inference for human body pose and shape estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 5253\u20105263.2020.","DOI":"10.1109\/CVPR42600.2020.00530"},{"key":"e_1_2_9_11_1","doi-asserted-by":"crossref","unstructured":"SongW WangX GaoY HaoA HouX.Real\u2010time expressive avatar animation generation based on monocular videos. 2022 IEEE international symposium on mixed and augmented reality adjunct (ISMAR\u2010adjunct). IEEE 429\u2010434.2022.","DOI":"10.1109\/ISMAR-Adjunct57072.2022.00092"},{"key":"e_1_2_9_12_1","doi-asserted-by":"crossref","unstructured":"LiW LiuH TangH WangP Van GoolL.Mhformer: multi\u2010hypothesis transformer for 3d human pose estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 13147\u201013156.2022.","DOI":"10.1109\/CVPR52688.2022.01280"},{"key":"e_1_2_9_13_1","doi-asserted-by":"crossref","unstructured":"ShanW LiuZ ZhangX WangS MaS GaoW.P\u2010stmo: pre\u2010trained spatial temporal many\u2010to\u2010one model for 3d human pose estimation. European conference on computer vision. Springer 461\u2010478.2022.","DOI":"10.1007\/978-3-031-20065-6_27"},{"key":"e_1_2_9_14_1","doi-asserted-by":"crossref","unstructured":"SunX XiaoB WeiF LiangS WeiY.Integral human pose regression. Proceedings of the European conference on computer vision (ECCV) 529\u2010545.2018.","DOI":"10.1007\/978-3-030-01231-1_33"},{"key":"e_1_2_9_15_1","doi-asserted-by":"crossref","unstructured":"ZhouK HanX JiangN JiaK LuJ.Hemlets pose: learning part\u2010centric heatmap triplets for accurate 3d human pose estimation. Proceedings of the IEEE\/CVF international conference on computer vision 2344\u20102353.2019.","DOI":"10.1109\/ICCV.2019.00243"},{"key":"e_1_2_9_16_1","doi-asserted-by":"crossref","unstructured":"ChengB XiaoB WangJ ShiH HuangTS ZhangL.Higherhrnet: scale\u2010aware representation learning for bottom\u2010up human pose estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 5386\u20105395.2020.","DOI":"10.1109\/CVPR42600.2020.00543"},{"key":"e_1_2_9_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3330016"},{"key":"e_1_2_9_18_1","doi-asserted-by":"crossref","unstructured":"YangS QuanZ NieM YangW.Transpose: Keypoint localization via transformer. Proceedings of the IEEE\/CVF international conference on computer vision 11802\u201011812.2021.","DOI":"10.1109\/ICCV48922.2021.01159"},{"key":"e_1_2_9_19_1","doi-asserted-by":"crossref","unstructured":"LiY ZhangS WangZ et al.Tokenpose: learning keypoint tokens for human pose estimation. Proceedings of the IEEE\/CVF. International Conference on Computer Vision 11313\u201011322.2021.","DOI":"10.1109\/ICCV48922.2021.01112"},{"key":"e_1_2_9_20_1","doi-asserted-by":"crossref","unstructured":"LuoZ WangZ HuangY WangL TanT ZhouE.Rethinking the heatmap regression for bottom\u2010up human pose estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 13264\u201013273.2021.","DOI":"10.1109\/CVPR46437.2021.01306"},{"key":"e_1_2_9_21_1","first-page":"30","article-title":"Associative embedding: end\u2010to\u2010end learning for joint detection and grouping","author":"Newell A","year":"2017","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_2_9_22_1","doi-asserted-by":"crossref","unstructured":"YangZ ZengA YuanC LiY.Effective whole\u2010body pose estimation with two\u2010stages distillation. Proceedings of the IEEE\/CVF. International Conference on Computer Vision 4210\u20104220.2023.","DOI":"10.1109\/ICCVW60793.2023.00455"},{"key":"e_1_2_9_23_1","doi-asserted-by":"crossref","unstructured":"WangY LiM CaiH ChenWM HanS.Lite pose: efficient architecture design for 2d human pose estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 13126\u201013136.2022.","DOI":"10.1109\/CVPR52688.2022.01278"},{"key":"e_1_2_9_24_1","unstructured":"VenkataramananS GhodratiA AsanoYM PorikliF HabibianA.Skip\u2010Attention: Improving Vision Transformers by Paying Less Attention. arXiv preprint arXiv:2301.022402023."},{"key":"e_1_2_9_25_1","unstructured":"KornblithS NorouziM LeeH HintonG.Similarity of neural network representations revisited. International conference on machine learning. Pmlr 3519\u20103529.2019."},{"key":"e_1_2_9_26_1","doi-asserted-by":"publisher","DOI":"10.1002\/0470045345"},{"key":"e_1_2_9_27_1","doi-asserted-by":"crossref","unstructured":"LinTY MaireM BelongieS et al.Microsoft coco: common objects in context. Computer vision\u2013ECCV 2014: 13th European conference Zurich Switzerland September 6\u201012 2014 proceedings part V 13. Springer 740\u2010755.2014.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_2_9_28_1","doi-asserted-by":"crossref","unstructured":"AndrilukaM PishchulinL GehlerP SchieleB.2d human pose estimation: new benchmark and state of the art analysis. Proceedings of the IEEE conference on computer vision and pattern recognition 3686\u20103693.2014.","DOI":"10.1109\/CVPR.2014.471"},{"key":"e_1_2_9_29_1","doi-asserted-by":"crossref","unstructured":"SunK XiaoB LiuD WangJ.Deep high\u2010resolution representation learning for human pose estimation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition 5693\u20105703.2019.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"e_1_2_9_30_1","unstructured":"YuanY RaoF LangH et al.Hrformer: High\u2010resolution transformer for dense prediction.2021arXiv preprint arXiv:2110.09408; 19."},{"key":"e_1_2_9_31_1","unstructured":"JiangT LuP ZhangL et al.RTMPose: Real\u2010Time Multi\u2010Person Pose Estimation based on MMPose. arXiv preprint arXiv:2303.073992023."},{"key":"e_1_2_9_32_1","unstructured":"YukihikoA.USB Camera Motion Capture ThreeDPoseTracker Description.2022."},{"key":"e_1_2_9_33_1","unstructured":"Live3D.VTuber.2023."},{"key":"e_1_2_9_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-023-03193-2"},{"key":"e_1_2_9_35_1","doi-asserted-by":"publisher","DOI":"10.1207\/s15327108ijap0303_3"}],"container-title":["Computer Animation and Virtual Worlds"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cav.2233","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T15:54:57Z","timestamp":1732118097000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/10.1002\/cav.2233"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5]]},"references-count":34,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,5]]}},"alternative-id":["10.1002\/cav.2233"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/cav.2233","archive":["Portico"],"relation":{},"ISSN":["1546-4261","1546-427X"],"issn-type":[{"value":"1546-4261","type":"print"},{"value":"1546-427X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5]]},"assertion":[{"value":"2024-04-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e2233"}}