{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T01:11:31Z","timestamp":1782436291725,"version":"3.54.5"},"reference-count":47,"publisher":"Institution of Engineering and Technology (IET)","issue":"1","license":[{"start":{"date-parts":[[2025,3,9]],"date-time":"2025-03-09T00:00:00Z","timestamp":1741478400000},"content-version":"vor","delay-in-days":67,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/creativecommons.org\/licenses\/by-nc\/4.0\/"},{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100010211","name":"Education Department of Jilin Province","doi-asserted-by":"publisher","award":["JJKH20220536KJ"],"award-info":[{"award-number":["JJKH20220536KJ"]}],"id":[{"id":"10.13039\/501100010211","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019594","name":"Yanbian University","doi-asserted-by":"publisher","award":["ydbq202204"],"award-info":[{"award-number":["ydbq202204"]}],"id":[{"id":"10.13039\/501100019594","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["ietresearch.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["IET Image Processing"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>In recent years, self\u2010supervised learning methods for monocular depth estimation have garnered significant attention due to their ability to learn from large amounts of unlabelled data. In this study, we propose further improvements for endoscopic scenes based on existing self\u2010supervised monocular depth estimation methods. The previous method introduce an appearance flow to address brightness inconsistencies caused by lighting changes and uses a unified self\u2010supervised framework to estimate both depth and camera motion simultaneously. However, to further enhance the model's supervisory signals, we introduce a new feature\u2010based perceptual loss. This module utilizes a pre\u2010trained encoder to extract features from both the synthesized and target frames and calculates their cosine dissimilarity as an additional source of supervision. In this way, we aim to improve the model's robustness in handling complex lighting and surface reflection conditions in endoscopic scenes. We compare the performance of using two pre\u2010trained CNN\u2010based models and four foundational models as encoder. Experimental results show that our improve method further enhances the accuracy of depth estimation in medical imaging. Additionally, it demonstrates that features extracted by CNN\u2010based models, which are sensitive to local details, outperform foundation models. This suggests that encoders for extracting medical image features may not require extensive pre\u2010training, and relatively simple traditional convolutional neural networks can\u00a0suffice.<\/jats:p>","DOI":"10.1049\/ipr2.70035","type":"journal-article","created":{"date-parts":[[2025,3,10]],"date-time":"2025-03-10T04:50:23Z","timestamp":1741582223000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Enhancing Self\u2010Supervised Monocular Depth Estimation in Endoscopy via Feature\u2010Based Perceptual Loss"],"prefix":"10.1049","volume":"19","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-4333-3538","authenticated-orcid":false,"given":"Kejin","family":"Zhu","sequence":"first","affiliation":[{"name":"College of Engineering Yanbian University Yanji Jilin Province China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Cui","sequence":"additional","affiliation":[{"name":"College of Engineering Yanbian University Yanji Jilin Province China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"265","published-online":{"date-parts":[[2025,3,9]]},"reference":[{"key":"e_1_2_11_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2021.102338"},{"key":"e_1_2_11_3_1","first-page":"2366","volume-title":"Advances in Neural Information Processing Systems","author":"Eigen D.","year":"2014"},{"key":"e_1_2_11_4_1","doi-asserted-by":"crossref","unstructured":"C.Liu J.Shen X.Hu L.Liu andF.Porikli \u201cLearning Data\u2010Driven Reflectance Priors for Intrinsic Image Decomposition \u201d inProceedings of the IEEE International Conference on Computer Vision(IEEE 2015) 3469\u20133477.","DOI":"10.1109\/ICCV.2015.396"},{"key":"e_1_2_11_5_1","unstructured":"W.Chen H.Fu Y.Yang Q.Deng X.Ding C.Tan et\u00a0al. \u201cSingle\u2010image Depth Perception in the Wild \u201d inProceedings of the IEEE Conference on Ccomputer Vision and Pattern Eecognition (IEEE 2016) 2713\u20132721."},{"key":"e_1_2_11_6_1","doi-asserted-by":"crossref","unstructured":"D.Xu E.Ricci W.Ouyang X.Wang andN.Sebe \u201cMulti\u2010Scale Continuous Crfs as Sequential Deep Networks for Monocular Depth Estimation \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 5354\u20135362.","DOI":"10.1109\/CVPR.2017.25"},{"key":"e_1_2_11_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2017.2740321"},{"key":"e_1_2_11_8_1","doi-asserted-by":"crossref","unstructured":"H.Fu M.Gong C.Wang K.Batmanghelich andD.Tao \u201cDeep Ordinal Regression Network for Monocular Depth Estimation \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 2002\u20132011.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"e_1_2_11_9_1","unstructured":"K.He A.Rafii andA.Godoy \u201cLearning Scene Structure and Depth from Monocular Images \u201darXiv:1803.07969(2018)."},{"key":"e_1_2_11_10_1","doi-asserted-by":"crossref","unstructured":"C.Xu S.Anwar andN.Barnes \u201cStructured Attention\u2010Guided Convolutional Neural Fields for Monocular Depth Estimation \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 2227\u20132236.","DOI":"10.1109\/CVPR.2018.00412"},{"key":"e_1_2_11_11_1","unstructured":"V.RepalaandS.Dubey \u201cDual\u2010path Multi\u2010Scale Fusion for Single Image Depth Estimation \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2019) 209\u2013217."},{"key":"e_1_2_11_12_1","unstructured":"Y.Shao R.Chen andR.Mahjourian \u201cSelf\u2010supervised Monocular Depth Estimation with Adaptive Geometric Consistency \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2021) 7160\u20137166."},{"key":"e_1_2_11_13_1","unstructured":"A.Geiger P.Lenz andR.Urtasun \u201cWe Present a New Dataset and Benchmark Suite for Visual Odometry and Monocular Depth Estimation \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2012) 335\u2013340."},{"key":"e_1_2_11_14_1","doi-asserted-by":"crossref","unstructured":"A.Saxena M.Sun andA. Y.Ng \u201cMake3D: Learning 3D Scene Structure from a Single Still Image \u201d inIEEE Transactions on Pattern Analysis and Machine Intelligence31 no.5(2008):824\u2013840.","DOI":"10.1109\/TPAMI.2008.132"},{"key":"e_1_2_11_15_1","unstructured":"I.Laina C.Rupprecht V.Belagiannis andT.Drummond \u201cDeeper Depth Prediction with Fully Convolutional Residual Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2016) 650\u2013658."},{"key":"e_1_2_11_16_1","doi-asserted-by":"crossref","unstructured":"A.RoyandS.Todorovic \u201cMonocular Depth Estimation Using Neural Regression Forest \u201d in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(IEEE 2016) 5506\u20135514.","DOI":"10.1109\/CVPR.2016.594"},{"key":"e_1_2_11_17_1","doi-asserted-by":"crossref","unstructured":"S.Shao Z.Pei X.Wu Z.Liu W.Chen andZ.Li \u201cIEBins: Iterative Elastic Bins for Monocular Depth Estimation \u201d inAdvances in Neural Information Processing Systems (NeurIPS)(ACM 2023) 53025\u201353037.","DOI":"10.52202\/075280-2307"},{"key":"e_1_2_11_18_1","doi-asserted-by":"crossref","unstructured":"S.Shao Z.Pei W.Chen X.Wu andZ.Li \u201cNddepth: Normal\u2010distance Assisted Monocular Depth Estimation \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2023) 7931\u20137940.","DOI":"10.1109\/ICCV51070.2023.00729"},{"key":"e_1_2_11_19_1","doi-asserted-by":"crossref","unstructured":"R.Ranftl A.Bochkovskiy andV.Koltun \u201cVision Transformers for Dense Prediction \u201d inProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(IEEE 2021) 12159\u201312168.","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_2_11_20_1","first-page":"6301","article-title":"Unsupervised Domain Adaptation for Depth Estimation via Conditional Generative Adversarial Networks","volume":"29","author":"Chen J.","year":"2020","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_2_11_21_1","doi-asserted-by":"crossref","unstructured":"J.Tremblay A.Prakash D.Acuna J.Gwak A.Agarwal C.Silva et\u00a0al. \u201cTraining Deep Networks with Synthetic Data: Bridging the Reality Gap by Domain Randomization \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)(IEEE 2018) 969\u2013977.","DOI":"10.1109\/CVPRW.2018.00143"},{"key":"e_1_2_11_22_1","doi-asserted-by":"crossref","unstructured":"M.Visentini\u2010Scarzanella X.Du J.Han andA.Handa \u201cA Deep Learning Approach to Monocular Endoscopic 3D Reconstruction using Synthetic Data \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 1089\u20131099.","DOI":"10.1007\/s11548-017-1609-2"},{"key":"e_1_2_11_23_1","doi-asserted-by":"crossref","unstructured":"F.MahmoodandN.Durr \u201cUnsupervised Monocular Depth Estimation with Synthetic Endoscopic Data \u201d inProceedings of the IEEE Conference on Medical Image Computing and Computer\u2010Assisted Intervention(IEEE 2018) 6602\u20136611.","DOI":"10.1109\/CVPR.2017.699"},{"key":"e_1_2_11_24_1","unstructured":"R.Chen F.Mahmood andN.Durr \u201cSelf\u2010Supervised Monocular Endoscopic Depth Estimation With Synthetic Data \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2019) 7063\u20137072."},{"key":"e_1_2_11_25_1","doi-asserted-by":"crossref","unstructured":"T.Zhou M.Brown N.Snavely andD. G.Lowe \u201cUnsupervised Learning of Depth and Ego\u2010Motion From Video \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 1851\u20131858.","DOI":"10.1109\/CVPR.2017.700"},{"key":"e_1_2_11_26_1","doi-asserted-by":"crossref","unstructured":"Z.Yang P.Wang Y.Wang W.Xu andR.Nevatia \u201cLEGO: Learning Edge With Geometry all at Once by Watching Videos \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (IEEE 2018) 225\u2013234.","DOI":"10.1109\/CVPR.2018.00031"},{"key":"e_1_2_11_27_1","doi-asserted-by":"crossref","unstructured":"R.Mahjourian M.Wicke andA.Angelova \u201cUnsupervised Learning of Depth and Ego\u2010Motion from Monocular Video Using 3D Geometric Constraints \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 5667\u20135675.","DOI":"10.1109\/CVPR.2018.00594"},{"key":"e_1_2_11_28_1","unstructured":"J.\u2010W.Bian Z.Li N.Wang H.Zhan C.Shen M.\u2010M.Cheng et\u00a0al. \u201cUnsupervised Scale\u2010Consistent Depth And Ego\u2010Motion Learning from Monocular Video \u201d inThirty\u2010Third Conference on Neural Information Processing Systems(ACM 2019) 35\u201345."},{"key":"e_1_2_11_29_1","doi-asserted-by":"crossref","unstructured":"H.Zhan R.Garg S.Weerasekera K.Li H.Agarwal andI.Reid \u201cUnsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2018) 340\u2013349.","DOI":"10.1109\/CVPR.2018.00043"},{"key":"e_1_2_11_30_1","doi-asserted-by":"crossref","unstructured":"J.Spencer R.Bowden andS.Hadfield \u201cDefeat\u2010Net: General Monocular Depth Via Simultaneous Unsupervised Representation Learning \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2020) 14\u00a0402\u201314\u00a0413.","DOI":"10.1109\/CVPR42600.2020.01441"},{"key":"e_1_2_11_31_1","unstructured":"X.Shu Z.Wang andB.Zhou \u201cFeature Fusion for Unsupervised Monocular Depth Estimation in Dynamic Scenes \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2020) 572\u2013588."},{"key":"e_1_2_11_32_1","unstructured":"T.Zhou S.Tulsiani N.Snavely andA. A.Efros \u201cUnsupervised Monocular Depth Estimation Through Self\u2010supervised Feature Learning \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2019 6872\u20136881."},{"key":"e_1_2_11_33_1","doi-asserted-by":"crossref","unstructured":"A.JohnstonandG.Carneiro \u201cSelf\u2010Supervised Monocular Trained Depth Estimation Using Self\u2010Attention and Discrete Disparity Volume \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2020) 4756\u20134765.","DOI":"10.1109\/CVPR42600.2020.00481"},{"key":"e_1_2_11_34_1","doi-asserted-by":"crossref","unstructured":"M.Turan E.Ornek N.Ibrahimli C.Giracoglu Y.Almalioglu M.Yanik et\u00a0al. \u201cUnsupervised Odometry and Depth Learning for Endoscopic Capsule Robots \u201d in2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)(IEEE 2018) 1801\u20131807.","DOI":"10.1109\/IROS.2018.8593623"},{"key":"e_1_2_11_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2021.102058"},{"key":"e_1_2_11_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2019.2950936"},{"key":"e_1_2_11_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3011067"},{"key":"e_1_2_11_38_1","unstructured":"M.Jaderberg K.Simonyan A.Zisserman et\u00a0al. \u201cSpatial Transformer Networks \u201d inAdvances in Neural Information Processing Systems (ACM 2015) 2017\u20132025."},{"key":"e_1_2_11_39_1","doi-asserted-by":"crossref","unstructured":"C.Godard O.Mac Aodha andG. J.Brostow \u201cUnsupervised Monocular Depth Estimation With Left\u2010Right Consistency \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 270\u2013279.","DOI":"10.1109\/CVPR.2017.699"},{"key":"e_1_2_11_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_2_11_41_1","doi-asserted-by":"crossref","unstructured":"K.He X.Zhang S.Ren andJ.Sun \u201cDeep Residual Learning for Image Recognition \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2016) 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_11_42_1","doi-asserted-by":"crossref","unstructured":"G.Huang Z.Liu L.van derMaaten andK. Q.Weinberger \u201cDensely Connected Convolutional Networks \u201d inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition(IEEE 2017) 4700\u20134708.","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_2_11_43_1","doi-asserted-by":"crossref","unstructured":"Z.Liu Y.Lin Y.Cao H.Hu Y.Wei Z.Zhang et\u00a0al. \u201cSwin transformer: Hierarchical Vision Transformer Using Shifted Windows \u201dProceedings of the IEEE\/CVF International Conference on Computer Vision(IEEE 2021) 10\u00a0012\u201310\u00a0022.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_2_11_44_1","unstructured":"A.Dosovitskiy L.Beyer A.Kolesnikov D.Weissenborn X.Zhai T.Unterthiner et\u00a0al. \u201cAn Image is Worth 16x16 Words: Transformers for Image Recognition at Scale \u201d inInternational Conference on Learning Representations(ACM 2021) 1\u201322."},{"key":"e_1_2_11_45_1","doi-asserted-by":"crossref","unstructured":"A.Kirillov E.Mintun N.Ravi H.Mao P.Rolland L.Gustafson et\u00a0al. \u201cSegment Anything \u201d inProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(IEEE 2023) 3992\u20134003.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_2_11_46_1","unstructured":"M.Oquab T.Darcet T.Moutakanni H.Vo M.Szafraniec V.Khalidov et\u00a0al. \u201cDINOv2: Learning Robust Visual Features Without Supervision \u201darXiv:2304.07193(2023)."},{"key":"e_1_2_11_47_1","unstructured":"M.Allan J.Mcleod C.Wang J.Rosenthal K.Fu T.Zeffiro et\u00a0al. \u201cStereo Correspondence and Reconstruction of Endoscopic Data Challenge \u201darXiv:2101.01133(2021)."},{"key":"e_1_2_11_48_1","unstructured":"A.Paszke S.Gross S.Chintala G.Chanan E.Yang Z.DeVito et\u00a0al. \u201cAutomatic differentiation in pytorch \u201d (2017) accessed July 10 2024 https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=BJJsrmfCZ."}],"container-title":["IET Image Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70035","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/ietresearch.onlinelibrary.wiley.com\/doi\/full-xml\/10.1049\/ipr2.70035","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/ietresearch.onlinelibrary.wiley.com\/doi\/pdf\/10.1049\/ipr2.70035","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T23:40:02Z","timestamp":1778283602000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/ietresearch.onlinelibrary.wiley.com\/doi\/10.1049\/ipr2.70035"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1]]}},"alternative-id":["10.1049\/ipr2.70035"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1049\/ipr2.70035","archive":["Portico"],"relation":{},"ISSN":["1751-9659","1751-9667"],"issn-type":[{"value":"1751-9659","type":"print"},{"value":"1751-9667","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2024-10-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-25","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70035"}}