{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:45:31Z","timestamp":1781534731751,"version":"3.54.5"},"reference-count":65,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T00:00:00Z","timestamp":1779062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["24-71-00112"],"award-info":[{"award-number":["24-71-00112"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]},{"award":["24-71-00112"],"award-info":[{"award-number":["24-71-00112"]}],"id":[{"id":"https:\/\/2.zoppoz.workers.dev:443\/https\/ror.org\/03y2gwe85","id-type":"ROR","asserted-by":"publisher"}]},{"name":"Russian state research","award":["FFZF-2025-0003"],"award-info":[{"award-number":["FFZF-2025-0003"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MTI"],"abstract":"<jats:p>Digital-avatar systems still provide limited control over emotionally expressive behavior in human\u2013computer interaction, especially in Large Language Model (LLM)-based chatbots and virtual assistants with personalized visual embodiments. To address this problem, we propose Multimodal Avatar Generation (MAVAGEN), a multimodal avatar generation framework for synthesizing upper-body digital avatars with personalized appearance and controllable emotional expression. The user specifies the desired gender and age, as well as provides a short text input from which the target emotional state is inferred. MAVAGEN then retrieves an identity image from the HaGRIDv2-1M corpus and generates an avatar clip with synchronized facial expressions, hand gestures, and expressive speech. The framework uses the following six feature streams: textual features, emotion-distribution features, landmark-based pose features, depth-geometry features, RGB-appearance features, and acoustic features. In a quantitative evaluation against recent human animation methods, MAVAGEN achieves the best overall avatar quality, with FID 48.20, FVD 592.00, SSIM 0.741, Sync-C 7.40, HKC 0.929, HKV 25.30, CSIM 0.563, and EmoAcc 0.88. Ablation results show that emotion and acoustic features contribute most to emotional agreement, while landmark-based pose and depth features improve geometric and motion stability. These results support the practical use of MAVAGEN in personalized LLM-based assistants and other emotion-sensitive interactive systems.<\/jats:p>","DOI":"10.3390\/mti10050055","type":"journal-article","created":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T16:07:26Z","timestamp":1779206846000},"page":"55","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MAVAGEN: Multimodal Avatar Generation Framework for Personalized Human\u2013Computer Interaction"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-7479-2851","authenticated-orcid":false,"given":"Alexandr","family":"Axyonov","sequence":"first","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-4135-6949","authenticated-orcid":false,"given":"Elena","family":"Ryumina","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-7935-0569","authenticated-orcid":false,"given":"Dmitry","family":"Ryumin","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-3424-652X","authenticated-orcid":false,"given":"Alexey","family":"Karpov","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gabriel, S., Puri, I., Xu, X., Malgaroli, M., and Ghassemi, M. (2024, January 12\u201316). Can AI Relate: Testing Large Language Model Response for Mental Health Support. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Miami, FL, USA.","DOI":"10.18653\/v1\/2024.findings-emnlp.120"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fei, H., Zhang, H., Wang, B., Liao, L., Liu, Q., and Cambria, E. (2024, January 11\u201316). EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-demos.7"},{"key":"ref_3","unstructured":"Zhang, H., Meng, Z., Luo, M., Han, H., Liao, L., Cambria, E., and Fei, H. (May, January 28). Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark. Proceedings of the ACM on Web Conference (WWW), Sydney, NSW, Australia."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1016\/j.neucom.2022.04.049","article-title":"Multitask Learning for Emotion and Personality Traits Detection","volume":"493","author":"Li","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wen, Z., Cao, J., Yang, Y., Yang, R., and Liu, S. (2024, January 11\u201315). Affective-NLI: Towards Accurate and Interpretable Personality Recognition in Conversation. Proceedings of the IEEE International Conference on Pervasive Computing and Communications (PerCom), Biarritz, France.","DOI":"10.1109\/PerCom59722.2024.10494487"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"122441","DOI":"10.1016\/j.eswa.2023.122441","article-title":"OCEAN-AI Framework with EmoFormer Cross-Hemiface Attention Approach for Personality Traits Assessment","volume":"239","author":"Ryumina","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Y., Xing, X., Lin, J., Zheng, H., Wang, Z., Liu, Q., and Xu, X. (2023, January 6\u201310). SoulChat: Improving LLMs\u2019 Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore.","DOI":"10.18653\/v1\/2023.findings-emnlp.83"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, Y., Yan, S., Liu, S., Li, Y., and Xiao, Y. (2024, January 11\u201316). EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.findings-acl.128"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Kyung, J., Heo, S., and Chang, J.H. (2024, January 1\u20135). Enhancing Multimodal Emotion Recognition through ASR Error Compensation and LLM Fine-Tuning. Proceedings of the Interspeech, Kos, Greece.","DOI":"10.21437\/Interspeech.2024-2364"},{"key":"ref_10","unstructured":"Xie, Y., Sun, C., Cao, Z., Liu, B., Ji, Z., Liu, Y., and Shan, L. (2025, January 9\u201314). A Dual Contrastive Learning Framework for Enhanced Multimodal Conversational Emotion Recognition. Proceedings of the International Conference on Computational Linguistics (COLING), Abu Dhabi, United Arab Emirates."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"110805","DOI":"10.52202\/079017-3518","article-title":"Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning","volume":"37","author":"Cheng","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst. (Neurips)"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xiao, M., Xie, Q., Kuang, Z., Liu, Z., Yang, K., Peng, M., Han, W., and Huang, J. (2024, January 11\u201316). HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.93"},{"key":"ref_13","unstructured":"Chiruzzo, L., Ritter, A., and Wang, L. (2025). Evaluating Vision-Language Models for Emotion Recognition. Findings of the Association for Computational Linguistics: NAACL 2025, Association for Computational Linguistics."},{"key":"ref_14","unstructured":"Wang, Y., Guo, J., Bai, J., Yu, R., He, T., Tan, X., Sun, X., and Bian, J. (March, January 25). Instructavatar: Text-guided emotion and motion control for avatar generation. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_15","unstructured":"Liu, T., Ma, Z., Chen, Q., Chen, F., Fan, S., Chen, X., and Yu, K. (March, January 25). VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_16","unstructured":"Song, W., Ding, Y., Hou, F., Li, S., Hao, A., and Hou, X. (March, January 25). CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible Networks. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, H., Sun, W., Di, D., Sun, S., Yang, J., Zou, C., and Bao, H. (2025, January 11\u201315). Moee: Mixture of emotion experts for audio-driven portrait animation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.02442"},{"key":"ref_18","unstructured":"Wei, X., Chen, P., Lu, M., Chen, H., and Tian, F. (March, January 25). Graphavatar: Compact head avatars with gnn-generated 3d gaussians. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Cha, H., Lee, I., and Joo, H. (2025, January 11\u201315). Perse: Personalized 3d generative avatars from a single portrait. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01487"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhang, D., Liu, Y., Lin, L., Zhu, Y., Chen, K., Qin, M., Li, Y., and Wang, H. (2025, January 11\u201315). HRAvatar: High-Quality and Relightable Gaussian Head Avatar. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.02448"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Feng, W.Q., Han, D., Zhou, Z.K., Li, S., Liu, X., Wan, P., Zhang, D., and Wang, M. (2025, January 11\u201315). GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projections. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.00032"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Corona, E., Zanfir, A., Bazavan, E.G., Kolotouros, N., Alldieck, T., and Sminchisescu, C. (2025, January 11\u201315). Vlogger: Multimodal diffusion for embodied avatar synthesis. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01482"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Qi, X., Pan, J., Li, P., Yuan, R., Chi, X., Li, M., Luo, W., Xue, W., Zhang, S., and Liu, Q. (2024, January 17\u201321). Weakly-supervised emotion transition learning for diverse 3d co-speech gesture generation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00992"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yariv, G., Gat, I., Benaim, S., Wolf, L., Schwartz, I., and Adi, Y. (2024, January 20\u201327). Diverse and aligned audio-to-video generation via text-to-video model adaptation. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i7.28486"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Qin, M., Liu, Y., Xu, Y., Zhao, X., Liu, Y., and Wang, H. (2024, January 20\u201327). High-fidelity 3d head avatars reconstruction through spatially-varying expression conditioned neural radiance field. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i5.28256"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Lin, W., Zheng, C., Yong, J.H., and Xu, F. (2024, January 20\u201327). Relightable and animatable neural avatars from videos. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i4.28136"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hu, J., Liu, Y., Zhao, J., and Jin, Q. (2021, January 11\u201316). MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2021.acl-long.440"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Joshi, A., Bhat, A., Jain, A., Singh, A., and Modi, A. (2022, January 10\u201315). COGMEN: COntextualized GNN based Multimodal Emotion recognitioN. Proceedings of the Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.naacl-main.306"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Li, D., Wang, Y., Funakoshi, K., and Okumura, M. (2023, January 6\u201310). Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion Recognition. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.996"},{"key":"ref_30","unstructured":"Yun, T., Lim, H., Lee, J., and Song, M. (2024, January 16\u201321). TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation. Proceedings of the Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), Mexico City, Mexico."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"4908","DOI":"10.1109\/TNNLS.2024.3367940","article-title":"DER-GCN: Dialog and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialog Emotion Recognition","volume":"36","author":"Ai","year":"2025","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"11886","DOI":"10.1109\/TCSVT.2024.3424777","article-title":"CEPrompt: Cross-Modal Emotion-Aware Prompting for Facial Expression Recognition","volume":"34","author":"Zhou","year":"2024","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_33","unstructured":"Liu, Y., Huang, Y., Liu, S., Zhan, Y., Chen, Z., and Chen, Z. (November, January 28). Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting. Proceedings of the ACM International Conference on Multimedia, Melbourne, VIC, Australia."},{"key":"ref_34","unstructured":"Murzaku, J., and Rambow, O. (2025). OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3069","DOI":"10.1109\/TMM.2025.3557704","article-title":"ExpLLM: Towards Chain of Thought for Facial Expression Recognition","volume":"27","author":"Lan","year":"2025","journal-title":"IEEE Trans. Multimed."},{"key":"ref_36","unstructured":"Wu, Z., Jiang, L., Li, X., Fang, C., Qin, Y., and Li, G. (March, January 25). Hierarchically controlled deformable 3D gaussians for talking head synthesis. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_37","unstructured":"Xu, Q., Yuan, S., Wei, Y., Wu, J., Wang, L., and Wu, C. (March, January 25). Multiple Feature Refining Network for Visual Emotion Distribution Learning. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Xiang, J., Gao, X., Guo, Y., and Zhang, J. (2024, January 17\u201321). Flashavatar: High-fidelity head avatar with efficient gaussian embedding. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00177"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Cai, H., Xiao, Y., Wang, X., Li, J., Guo, Y., Fan, Y., Gao, S., and Zhang, J. (2025, January 11\u201315). HERA: Hybrid Explicit Representation for Ultra-Realistic Head Avatars. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.00033"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhan, Y., Shao, T., Yang, Y., and Zhou, K. (2025, January 11\u201315). Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.02449"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Ma, F., Fan, H., and Chua, T.S. (2025, January 11\u201315). Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01486"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Zhuang, J., Kang, D., Bao, L., Lin, L., and Li, G. (2025, January 11\u201315). Dagsm: Disentangled avatar generation with gs-enhanced mesh. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.00036"},{"key":"ref_43","unstructured":"Hu, L. (2024, January 17\u201321). Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Meng, R., Zhang, X., Li, Y., and Ma, C. (2025, January 11\u201315). EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.00516"},{"key":"ref_45","unstructured":"Kapitanov, A., Kvanchiani, K., Nagaev, A., Kraynov, R., and Makhliarchuk, A. (2024, January 4\u20138). HaGRID\u2013HAnd Gesture Recognition Image Dataset. Proceedings of the Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA."},{"key":"ref_46","unstructured":"Bazarevsky, V., Kartynnik, Y., Vakunov, A., Raveendran, K., and Grundmann, M. (2019, January 16\u201320). BlazeFace: Sub-Millisecond Neural Face Detection on Mobile GPUs. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA."},{"key":"ref_47","unstructured":"Kartynnik, Y., Ablavatski, A., Grishchenko, I., and Grundmann, M. (2019, January 16\u201320). Real-Time Facial Surface Geometry from Monocular Video on Mobile GPUs. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA."},{"key":"ref_48","unstructured":"Zhang, F., Bazarevsky, V., Vakunov, A., Tkachenka, A., Sung, G., Chang, C.L., and Grundmann, M. (2020, January 14\u201319). MediaPipe Hands: On-Device Real-Time Hand Tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA."},{"key":"ref_49","unstructured":"Bazarevsky, V., Grishchenko, I., Raveendran, K., Zhu, T., Zhang, F., and Grundmann, M. (2020, January 14\u201319). BlazePose: On-Device Real-Time Body Pose Tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA."},{"key":"ref_50","unstructured":"Dao, T., and Gu, A. (2024, January 21\u201327). Transformers are SSMs: Generalized Models and Efficient Algorithms through Structured State Space Duality. Proceedings of the International Conference on Machine Learning (ICML), Vienna, Austria. Available online: https:\/\/2.zoppoz.workers.dev:443\/https\/proceedings.mlr.press\/v235\/dao24a.html."},{"key":"ref_51","first-page":"2366","article-title":"Depth map prediction from a single image using a multi-scale deep network","volume":"27","author":"Eigen","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst. (NeurIPS)"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ke, B., Qu, K., Wang, T., Metzger, N., Huang, S., Li, B., Obukhov, A., and Schindler, K. (2025). Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. IEEE Trans. Pattern Anal. Mach. Intell., 1\u201318.","DOI":"10.1109\/TPAMI.2025.3591076"},{"key":"ref_53","unstructured":"Kingma, D.P., and Welling, M. (2014, January 14\u201316). Auto-Encoding Variational Bayes. Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Boncelet, C. (2009). Image noise models. The Essential Guide to Image Processing, Academic Press.","DOI":"10.1016\/B978-0-12-374457-9.00007-X"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Peng, Y., Sudo, Y., Shakeel, M., and Watanabe, S. (2024, January 11\u201316). OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.549"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Tan, C., Gao, Z., Wu, L., Xu, Y., Xia, J., Li, S., and Li, S.Z. (2023, January 18\u201322). Temporal attention unit: Towards efficient spatiotemporal predictive learning. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01800"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer International Publishing.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Huang, Z., Tang, F., Zhang, Y., Cun, X., Cao, J., Li, J., and Lee, T.Y. (2024, January 17\u201321). Make-your-anchor: A diffusion-based 2d avatar generation framework. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00668"},{"key":"ref_59","unstructured":"Loshchilov, I., and Hutter, F. (2019, January 6\u20139). Decoupled Weight Decay Regularization. Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA."},{"key":"ref_60","unstructured":"Loshchilov, I., and Hutter, F. (2017, January 24\u201326). SGDR: Stochastic Gradient Descent with Warm Restarts. Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France."},{"key":"ref_61","unstructured":"Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. (2018). Towards accurate generative models of video: A new metric & challenges. arXiv."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Prajwal, K., Mukhopadhyay, R., Namboodiri, V.P., and Jawahar, C. (2020, January 12\u201316). A lip sync expert is all you need for speech to lip generation in the wild. Proceedings of the ACM international conference on multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413532"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"104207","DOI":"10.1016\/j.inffus.2026.104207","article-title":"Multi-Lingual Approach for Multi-Modal Emotion and Sentiment Recognition Based on Triple Fusion","volume":"132","author":"Markitantov","year":"2026","journal-title":"Inf. Fusion"},{"key":"ref_65","unstructured":"Zhang, Y., Gu, J., Wang, L.W., Wang, H., Cheng, J., Zhu, Y., and Zou, F. (2025, January 13\u201319). MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance. Proceedings of the International Conference on Machine Learning (ICML), Vancouver, BC, Canada."}],"container-title":["Multimodal Technologies and Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.mdpi.com\/2414-4088\/10\/5\/55\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T16:40:17Z","timestamp":1779208817000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.mdpi.com\/2414-4088\/10\/5\/55"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":65,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["mti10050055"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.3390\/mti10050055","relation":{},"ISSN":["2414-4088"],"issn-type":[{"value":"2414-4088","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}