{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T05:42:15Z","timestamp":1757310135337,"version":"3.37.3"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2023,5,23]],"date-time":"2023-05-23T00:00:00Z","timestamp":1684800000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,5,23]],"date-time":"2023-05-23T00:00:00Z","timestamp":1684800000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["ZN 01IS17050"],"award-info":[{"award-number":["ZN 01IS17050"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J CARS"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>\n                           <jats:bold>Purpose<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>Image-to-image translation methods can address the lack of diversity in publicly available cataract surgery data. However, applying image-to-image translation to videos\u2014which are frequently used in medical downstream applications\u2014induces artifacts. Additional spatio-temporal constraints are needed to produce realistic translations and improve the temporal consistency of translated image sequences.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Methods<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>We introduce a motion-translation module that translates optical flows between domains to impose such constraints. We combine it with a shared latent space translation model to improve image quality. Evaluations are conducted regarding translated sequences\u2019 image quality and temporal consistency, where we propose novel quantitative metrics for the latter. Finally, the downstream task of surgical phase classification is evaluated when retraining it with additional synthetic translated data.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Results<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>Our proposed method produces more consistent translations than state-of-the-art baselines. Moreover, it stays competitive in terms of the per-image translation quality. We further show the benefit of consistently translated cataract surgery sequences for improving the downstream task of surgical phase prediction.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>\n                           <jats:bold>Conclusion<\/jats:bold>\n                        <\/jats:title>\n                <jats:p>The proposed module increases the temporal consistency of translated sequences. Furthermore, imposed temporal constraints increase the usability of translated data in downstream tasks. This allows overcoming some of the hurdles of surgical data acquisition and annotation and enables improving models\u2019 performance by translating between existing datasets of sequential frames.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1007\/s11548-023-02925-y","type":"journal-article","created":{"date-parts":[[2023,5,23]],"date-time":"2023-05-23T12:02:19Z","timestamp":1684843339000},"page":"1217-1224","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Temporally consistent sequence-to-sequence translation of cataract surgeries"],"prefix":"10.1007","volume":"18","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0005-8097-0158","authenticated-orcid":false,"given":"Yannik","family":"Frisch","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Moritz","family":"Fuchs","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anirban","family":"Mukhopadhyay","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,5,23]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"Wang W, Yan W, Fotis K, Prasad NM, Lansingh VC, Taylor HR, Finger RP, Facciolo D, He M (2016) Cataract surgical rate and socioeconomics: a global study. IOVS 57(14):5872\u20135881","key":"2925_CR1","DOI":"10.1167\/iovs.16-19894"},{"key":"2925_CR2","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1016\/j.media.2018.11.008","volume":"52","author":"H Al Hajj","year":"2019","unstructured":"Al Hajj H, Lamard M, Conze P-H, Roychowdhury S, Hu X, Mar\u0161alkait\u0117 G, Zisimopoulos O, Dedmari MA, Zhao F, Prellberg J et al (2019) Cataracts: challenge on automatic tool annotation for cataract surgery. Med Image Anal 52:24\u201341","journal-title":"Med Image Anal"},{"doi-asserted-by":"crossref","unstructured":"Zisimopoulos O, Flouty E, Luengo I, Giataganas P, Nehme J, Chow A, Stoyanov D (2018) Deepphase: surgical phase recognition in cataracts videos. In: MICCAI. Springer, pp 265\u2013272","key":"2925_CR3","DOI":"10.1007\/978-3-030-00937-3_31"},{"unstructured":"Luengo I, Grammatikopoulou M, Mohammadi R, Walsh C, Nwoye CI, Alapatt D, Padoy N, Ni Z-L, Fan C-C, Bian G-B et al. (2021) 2020 cataracts semantic segmentation challenge. arXiv:2110.10965","key":"2925_CR4"},{"doi-asserted-by":"crossref","unstructured":"Schoeffmann K, Taschwer M, Sarny S, M\u00fcnzer B, Primus MJ, Putzgruber D (2018) Cataract-101: video dataset of 101 cataract surgeries. In: ACMMMSYS, pp 421\u2013425","key":"2925_CR5","DOI":"10.1145\/3204949.3208137"},{"doi-asserted-by":"crossref","unstructured":"Zhu J-Y, Park T, Isola P, Efros AA (2017) Unpaired image-to-image translation using cycle-consistent adversarial networks. In: ECCV, pp 2223\u20132232","key":"2925_CR6","DOI":"10.1109\/ICCV.2017.244"},{"unstructured":"Liu M-Y, Breuel T, Kautz J (2017) Unsupervised image-to-image translation networks. In: NEURIPS, vol 30","key":"2925_CR7"},{"doi-asserted-by":"crossref","unstructured":"Bansal A, Ma S, Ramanan D, Sheikh Y (2018) Recycle-gan: unsupervised video retargeting. In: ECCV, pp 119\u2013135","key":"2925_CR8","DOI":"10.1007\/978-3-030-01228-1_8"},{"doi-asserted-by":"crossref","unstructured":"Liu K, Gu S, Romero A, Timofte R (2021) Unsupervised multimodal video-to-video translation via self-supervised learning. In: WACV, pp 1030\u20131040","key":"2925_CR9","DOI":"10.1109\/WACV48630.2021.00107"},{"doi-asserted-by":"crossref","unstructured":"Chen Y, Pan Y, Yao T, Tian X, Mei T (2019) Mocycle-gan: unpaired video-to-video translation. In: ACMMM, pp 647\u2013655","key":"2925_CR10","DOI":"10.1145\/3343031.3350937"},{"doi-asserted-by":"crossref","unstructured":"Huang X, Liu M-Y, Belongie S, Kautz J (2018) Multimodal unsupervised image-to-image translation. In: ECCV, pp 172\u2013189","key":"2925_CR11","DOI":"10.1007\/978-3-030-01219-9_11"},{"doi-asserted-by":"crossref","unstructured":"Lee H-Y, Tseng H-Y, Huang J-B, Singh M, Yang M-H (2018) Diverse image-to-image translation via disentangled representations. In: ECCV, pp 35\u201351","key":"2925_CR12","DOI":"10.1007\/978-3-030-01246-5_3"},{"doi-asserted-by":"crossref","unstructured":"You A, Kim JK, Ryu IH, Yoo TK (2022) Application of generative adversarial networks (GAN) for ophthalmology image domains: a survey. Eye Vis 9(1):1\u201319","key":"2925_CR13","DOI":"10.1186\/s40662-022-00277-3"},{"unstructured":"Skandarani Y, Jodoin P-M, Lalande A (2021) Gans for medical image synthesis: an empirical study. arXiv:2105.05318","key":"2925_CR14"},{"doi-asserted-by":"crossref","unstructured":"Shin H-C, Tenenholtz NA, Rogers JK, Schwarz CG, Senjem ML, Gunter JL, Andriole KP, Michalski M (2018) Medical image synthesis for data augmentation and anonymization using generative adversarial networks. In: SASHIMI. Springer, pp 1\u201311","key":"2925_CR15","DOI":"10.1007\/978-3-030-00536-8_1"},{"issue":"6","key":"2925_CR16","doi-asserted-by":"publisher","first-page":"493","DOI":"10.1038\/s41551-021-00751-8","volume":"5","author":"RJ Chen","year":"2021","unstructured":"Chen RJ, Lu MY, Chen TY, Williamson DF, Mahmood F (2021) Synthetic data in machine learning for medicine and healthcare. Nat Biomed Eng 5(6):493\u2013497","journal-title":"Nat Biomed Eng"},{"doi-asserted-by":"crossref","unstructured":"Armanious K, Jiang C, Abdulatif S, K\u00fcstner T, Gatidis S, Yang B (2019) Unsupervised medical image translation using cycle-medgan. In: EUSIPCO. IEEE, pp 1\u20135","key":"2925_CR17","DOI":"10.23919\/EUSIPCO.2019.8902799"},{"key":"2925_CR18","first-page":"1964","volume":"34","author":"L Kong","year":"2021","unstructured":"Kong L, Lian C, Huang D, Hu Y, Zhou Q et al (2021) Breaking the dilemma of medical image-to-image translation. NEURIPS 34:1964\u20131978","journal-title":"NEURIPS"},{"doi-asserted-by":"crossref","unstructured":"Pfeiffer M, Funke I, Robu MR, Bodenstedt S, Strenger L, Engelhardt S, Ro\u00df T, Clarkson MJ, Gurusamy K, Davidson BR et al (2019) Generating large labeled data sets for laparoscopic image processing tasks using unpaired image-to-image translation. In: MICCAI. Springer, pp 119\u2013127","key":"2925_CR19","DOI":"10.1007\/978-3-030-32254-0_14"},{"doi-asserted-by":"crossref","unstructured":"Rivoir D, Pfeiffer M, Docea R, Kolbinger F, Riediger C, Weitz J, Speidel S (2021) Long-term temporally consistent unpaired video translation from simulated surgical 3d data. In: ICCV, pp 3343\u20133353","key":"2925_CR20","DOI":"10.1109\/ICCV48922.2021.00333"},{"doi-asserted-by":"crossref","unstructured":"Sahu M, Mukhopadhyay A, Zachow S (2021) Simulation-to-real domain adaptation with teacher-student learning for endoscopic instrument segmentation. IJCARS 16(5):849\u2013859","key":"2925_CR21","DOI":"10.1007\/s11548-021-02383-4"},{"doi-asserted-by":"crossref","unstructured":"Park K, Woo S, Kim D, Cho D, Kweon IS (2019) Preserving semantic and temporal consistency for unpaired video-to-video translation. In: ACMMM, pp 1248\u20131257","key":"2925_CR22","DOI":"10.1145\/3343031.3350864"},{"key":"2925_CR23","first-page":"1083","volume":"33","author":"C Lei","year":"2020","unstructured":"Lei C, Xing Y, Chen Q (2020) Blind video temporal consistency via deep video prior. NEURIPS 33:1083\u20131093","journal-title":"NEURIPS"},{"doi-asserted-by":"crossref","unstructured":"Teed Z, Deng J (2020) Raft: recurrent all-pairs field transforms for optical flow. In: ECCV. Springer, pp 402\u2013419","key":"2925_CR24","DOI":"10.1007\/978-3-030-58536-5_24"},{"doi-asserted-by":"crossref","unstructured":"Ronneberger O, Fischer P, Brox T (2015) U-net: convolutional networks for biomedical image segmentation. In: MICCAI. Springer, pp 234\u2013241","key":"2925_CR25","DOI":"10.1007\/978-3-319-24574-4_28"},{"doi-asserted-by":"crossref","unstructured":"Zhang R, Isola P, Efros AA, Shechtman E, Wang O (2018) The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR, pp 586\u2013595","key":"2925_CR26","DOI":"10.1109\/CVPR.2018.00068"},{"doi-asserted-by":"crossref","unstructured":"Chu M, Xie Y, Mayer J, Leal-Taix\u00e9 L, Thuerey N (2020) Learning temporal coherence via self-supervision for GAN-based video generation. ACMTOG 39(4):75-1","key":"2925_CR27","DOI":"10.1145\/3386569.3392457"}],"container-title":["International Journal of Computer Assisted Radiology and Surgery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/content\/pdf\/10.1007\/s11548-023-02925-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/article\/10.1007\/s11548-023-02925-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/content\/pdf\/10.1007\/s11548-023-02925-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,8]],"date-time":"2023-07-08T16:09:10Z","timestamp":1688832550000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/10.1007\/s11548-023-02925-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,23]]},"references-count":27,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["2925"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1007\/s11548-023-02925-y","relation":{},"ISSN":["1861-6429"],"issn-type":[{"type":"electronic","value":"1861-6429"}],"subject":[],"published":{"date-parts":[[2023,5,23]]},"assertion":[{"value":"14 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 April 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 May 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}