{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:53:28Z","timestamp":1777658008585,"version":"3.51.4"},"reference-count":79,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T00:00:00Z","timestamp":1716163200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T00:00:00Z","timestamp":1716163200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"University of the Witwatersrand"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The restricted access to data in healthcare facilities due to patient privacy and confidentiality policies has led to the application of general natural language processing (NLP) techniques advancing relatively slowly in the health domain. Additionally, because clinical data is unique to various institutions and laboratories, there are not enough standards and conventions for data annotation. In places without robust death registration systems, the cause of death (COD) is determined through a verbal autopsy (VA) report. A non-clinician field agent completes a VA report using a set of standardized questions as guide to identify the symptoms of a COD. The narrative text of the VA report is used as a case study to examine the difficulties of applying NLP techniques to the healthcare domain. This paper presents a framework that leverages knowledge across multiple domains via two domain adaptation techniques: feature extraction and fine-tuning. These techniques aim to improve VA text representations for COD classification tasks in the health domain. The framework is motivated by multi-step learning, where a final learning task is realized via a sequence of intermediate learning tasks. The framework builds upon the strengths of the Bidirectional Encoder Representations from Transformers (BERT) and Embeddings from Language Models (ELMo) models pretrained on the general English and biomedical domains. These models are employed to extract features from the VA narratives. Our results demonstrate improved performance when initializing the learning of BERT embeddings with ELMo embeddings. The benefit of incorporating character-level information for learning word embeddings in the English domain, coupled with word-level information for learning word embeddings in the biomedical domain, is also evident.<\/jats:p>","DOI":"10.1007\/s11063-024-11526-y","type":"journal-article","created":{"date-parts":[[2024,5,20]],"date-time":"2024-05-20T08:01:42Z","timestamp":1716192102000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Multi-step Transfer Learning in Natural Language Processing for the Health Domain"],"prefix":"10.1007","volume":"56","author":[{"given":"Thokozile","family":"Manaka","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Terence Van","family":"Zyl","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deepak","family":"Kar","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alisha","family":"Wade","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,5,20]]},"reference":[{"key":"11526_CR1","unstructured":"United Nations (2013) Department of economic and social affairs, population division, united nations. World Population Prospects: The 2012 revision"},{"key":"11526_CR2","volume-title":"Verbal autopsy standards: ascertaining and attributing cause of death, Geneva","author":"World Health Organisation","year":"2007","unstructured":"World Health Organisation (2007) Verbal autopsy standards: ascertaining and attributing cause of death, Geneva. Switzerland, World Health Organisation"},{"issue":"5","key":"11526_CR3","first-page":"450","volume":"18","author":"L Hirschman","year":"2011","unstructured":"Hirschman L, Chapman WW, D\u2019Avolio LW, Savova GK, Uzuner O (2011) Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions. J Am Med Inform Assoc 18(5):450\u2013453","journal-title":"J Am Med Inform Assoc"},{"key":"11526_CR4","doi-asserted-by":"publisher","first-page":"544","DOI":"10.1136\/amiajnl-2011-000464","volume":"18","author":"L Ohno-Machado","year":"2011","unstructured":"Ohno-Machado L, Nadkarni P, Chapman W (2011) Natural language processing: an introduction. J Am Med Inform Assoc 18:544\u201351","journal-title":"J Am Med Inform Assoc"},{"issue":"10","key":"11526_CR5","doi-asserted-by":"publisher","first-page":"1345","DOI":"10.1109\/TKDE.2009.191","volume":"22","author":"SJ Pan","year":"2010","unstructured":"Pan SJ, Yang Q (2010) A survey on transfer learning. IEEE Trans Knowl Data Eng 22(10):1345\u20131359","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"11526_CR6","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805"},{"issue":"8","key":"11526_CR7","doi-asserted-by":"publisher","first-page":"1202","DOI":"10.3390\/electronics11081202","volume":"11","author":"N Kooverjee","year":"2022","unstructured":"Kooverjee N, James S, Van Zyl T (2022) Investigating transfer learning in graph neural networks. Electronics 11(8):1202","journal-title":"Electronics"},{"key":"11526_CR8","doi-asserted-by":"crossref","unstructured":"Bhana N, van Zyl TL (2022) Knowledge graph fusion for language model fine-tuning. In: 2022 9th international conference on soft computing and machine intelligence (ISCMI)","DOI":"10.1109\/ISCMI56532.2022.10068451"},{"key":"11526_CR9","doi-asserted-by":"crossref","unstructured":"Kim Y (2014) Convolutional neural networks for sentence classification. In: Proceedings of the conference on empirical methods in natural language processing, pp 1746\u20131751","DOI":"10.3115\/v1\/D14-1181"},{"key":"11526_CR10","doi-asserted-by":"crossref","unstructured":"Ramachandran P, Liu PJ, Le QV (2016) Unsupervised pretraining for sequence to sequence learning. arXiv:1611.02683","DOI":"10.18653\/v1\/D17-1039"},{"key":"11526_CR11","doi-asserted-by":"crossref","unstructured":"Delrue, L., Gosselin, R., Ilsen, B., Landeghem, A.V., de Mey, J., Duyck, P.: Difficulties in the interpretation of chest radiography. Comparative Interpretation of CT and Standard Radiography of the Chest, 27\u201349 (2011)","DOI":"10.1007\/978-3-540-79942-9_2"},{"issue":"1","key":"11526_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1111\/1754-9485.12014","volume":"57","author":"SK Goergen","year":"2013","unstructured":"Goergen SK, Pool FJ, Turner TJ, Grimm JE, Appleyard MN, Crock C, Fahey MC, Fay MF, Ferris NJ, Liew SM, Perry RD, Revell A, Russell GM, Wang SC, Wriedt C (2013) Evidence-based guideline for the written radiology report: methods, recommendations and implementation challenges. J Med Imaging Radiat Oncol 57(1):1\u20137","journal-title":"J Med Imaging Radiat Oncol"},{"key":"11526_CR13","first-page":"3","volume":"81","author":"A Brady","year":"2012","unstructured":"Brady A, Laoide R, Mccarthy P, Mcdermott R (2012) Discrepancy and error in radiology: concepts, causes and consequences. Ulster Med J 81:3\u20139","journal-title":"Ulster Med J"},{"key":"11526_CR14","unstructured":"Liu F, You C, Wu X, Ge S, Sun X (2021) Auto-encoding knowledge graph for unsupervised medical report generation. CoRR abs\/2111.04318"},{"key":"11526_CR15","first-page":"18864","volume":"35","author":"F Liu","year":"2022","unstructured":"Liu F, Yang B, You C, Wu X, Ge S, Liu Z, Sun X, Yang Y, Clifton D (2022) Retrieve, reason, and refine: generating accurate and faithful patient instructions. NeurIPS 35:18864\u201318877","journal-title":"NeurIPS"},{"key":"11526_CR16","unstructured":"Li J, Wang X, Wu X, Zhang Z, Xu X, Fu J, Tiwari P, Wan X, Wang B (2023) Huatuo-26m, a large-scale chinese medical qa dataset. CoRR abs\/2305.01526"},{"key":"11526_CR17","unstructured":"Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, Steinhardt J (2020) Measuring massive multitask language understanding. CoRR abs\/2009.03300"},{"key":"11526_CR18","doi-asserted-by":"crossref","unstructured":"Abacha AB, Shivade C, Demner-Fushman D (2019) Overview of the mediqa 2019 shared task on textual inference, question entailment and question answering. In: Proceedings of the 18th BioNLP Workshop and Shared Task, pp 370\u2013379","DOI":"10.18653\/v1\/W19-5039"},{"key":"11526_CR19","first-page":"21916","volume":"35","author":"P Zhou","year":"2022","unstructured":"Zhou P, Wang Z, Chong D, Guo Z, Hua Y, Su Z, Teng Z, Wu J, Yang J (2022) Mets-cov: A dataset of medical entity and targeted sentiment on covid-19 related tweets. NeurIPS 35:21916\u201321932","journal-title":"NeurIPS"},{"key":"11526_CR20","unstructured":"Nori H, King N, McKinney SM, Carignan D, Horvitz E (2023) Capabilities of gpt-4 on medical challenge problems. CoRR abs\/2303.13375"},{"key":"11526_CR21","doi-asserted-by":"crossref","unstructured":"Fang C, Ling J, Zhou J, Wang Y, Liu X, Jiang Y, Wu Y, Chen Y, Zhu Z, Ma J, Yan Z (2023) How does chatgpt4 preform on non-english national medical licensing examination? an evaluation in chinese language. medRxiv 35","DOI":"10.1101\/2023.05.03.23289443"},{"key":"11526_CR22","doi-asserted-by":"crossref","unstructured":"Zeng Q, Garay L, Zhou P, Chong D, Hua Y, Wu J, Pan Y, Zhou H, Voigt R, Yang J (2022) Greenplm: Cross-lingual transfer of monolingual pre-trained language models at almost no cost. The 32nd International Joint Conference on Artificial Intelligence","DOI":"10.24963\/ijcai.2023\/698"},{"key":"11526_CR23","doi-asserted-by":"crossref","unstructured":"Liu J, Zhou P, Hua Y, Chong D, Tian Z, Liu A, Wang H, You C, Guo Z, Zhu L, Li M (2023) Benchmarking large language models on cmexam - a comprehensive chinese medical exam dataset. CoRR abs\/2306.03030","DOI":"10.1101\/2024.04.24.24306315"},{"key":"11526_CR24","doi-asserted-by":"crossref","unstructured":"Liu F, Zhu T, Wu X, Yang B, You C, Wang C, Lu L, Liu Z, Zheng Y, Sun X, Yang Y, Clifton L, Clifton DA (2023) A medical multimodal large language model for future pandemics. npj Digit. Med 6:226","DOI":"10.1038\/s41746-023-00952-2"},{"key":"11526_CR25","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1613\/jair.731","volume":"12","author":"J Baxter","year":"2000","unstructured":"Baxter J (2000) A model of inductive bias learning. J Artific Intell Res 12:149\u2013198","journal-title":"J Artific Intell Res"},{"key":"11526_CR26","doi-asserted-by":"crossref","unstructured":"Huang Z, Zweig G, Dmoulin B (2014) Cache based recurrent neural network language model inference for first pass speech recognition. IEEE ICASSP, pp 6354\u20136358","DOI":"10.1109\/ICASSP.2014.6854827"},{"key":"11526_CR27","doi-asserted-by":"crossref","unstructured":"Wen Z, Lu X, Reddy S (2020) Medal: Medical abbreviation disambiguation dataset for natural language understanding pretraining. Proceedings of the 3rd clinical natural language processing workshop, pp 130\u2013135","DOI":"10.18653\/v1\/2020.clinicalnlp-1.15"},{"key":"11526_CR28","doi-asserted-by":"crossref","unstructured":"Alsentzer E, Murphy JR, Boag W, Weng WH, Jin D, Naumann T, McDermott MBA (2019) Publicly available clinical bert embeddings. arXiv:1904.03323","DOI":"10.18653\/v1\/W19-1909"},{"issue":"4","key":"11526_CR29","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","volume":"36","author":"J Lee","year":"2020","unstructured":"Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, Kang J (2020) Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4):1234\u20131240","journal-title":"Bioinformatics"},{"key":"11526_CR30","unstructured":"Qiao J, Bhuwan D, William C, Xinghua L (2019) Probing biomedical embeddings from language models. In: Proceedings of the 3rd workshop on evaluating vector space representations for NLP, pp 82\u201389"},{"key":"11526_CR31","unstructured":"Beltagy I, Cohan A, Lo K (2019) Scibert: pretrained contextualized embeddings for scientific text. arXiv:1903.10676"},{"key":"11526_CR32","doi-asserted-by":"crossref","unstructured":"Johnson AEW, Pollard TJ, Shen L, Lehman LH, Feng M, Ghassemi M, Moody B, Szolovits P, Celi LA, Mark RG (2016) Mimic-III, a freely accessible critical care database. Sci Data3","DOI":"10.1038\/sdata.2016.35"},{"key":"11526_CR33","doi-asserted-by":"crossref","unstructured":"Peters M, Ruder S, Smith N (2019) To tune or not to tune? adapting pretrained representations to diverse tasks. arXiv:1903.05987","DOI":"10.18653\/v1\/W19-4302"},{"key":"11526_CR34","doi-asserted-by":"crossref","unstructured":"Jin Q, Dhingra B, Cohen W, Lu X (2019) Probing biomedical embeddings from language models. arXiv:1904.02181","DOI":"10.18653\/v1\/W19-2011"},{"key":"11526_CR35","unstructured":"Zhao S, Li B, Reed C, Xu P, Keutzer K (2020) Multi-source domain adaptation in the deep learning era: a systematic survey. arXiv:2002.12169"},{"key":"11526_CR36","doi-asserted-by":"crossref","unstructured":"Torralba A, Efros AA (2011) Unbiased look at dataset bias. In CVPR","DOI":"10.1109\/CVPR.2011.5995347"},{"key":"11526_CR37","doi-asserted-by":"crossref","unstructured":"Zhao S, Zhao X, Ding G, Keutzer K (2018) Emotiongan: Un-supervised domain adaptation for learning discrete probability distributions of image emotions. In ACM MM","DOI":"10.1145\/3240508.3240591"},{"key":"11526_CR38","unstructured":"III HD (2007) Frustratingly easy domain adaptation. Association for Computational Linguistic (ACL), pp 256\u2013263"},{"key":"11526_CR39","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1016\/j.inffus.2014.12.003","volume":"24","author":"S Sun","year":"2015","unstructured":"Sun S, Shi H, Wu Y (2015) A survey of multi-source domain adaptation. Inf Fusion 24:84\u201392","journal-title":"Inf Fusion"},{"key":"11526_CR40","unstructured":"Riemer M, Cases I, Ajemian R, Liu M, Rish I, Tu Y, Tesauro G (2019) Learning to learn without forgetting by maximizing transfer and minimizing interference. In ICLR"},{"key":"11526_CR41","first-page":"505","volume":"24","author":"Q Sun","year":"2011","unstructured":"Sun Q, Chattopadhyay R, Panchanathan S, Ye J (2011) A two-stage weighting framework for multi-source domain adaptation. Adv Neural Inform Process Syst 24:505\u2013513","journal-title":"Adv Neural Inform Process Syst"},{"key":"11526_CR42","first-page":"1433","volume":"21","author":"G Schweikert","year":"2009","unstructured":"Schweikert G, R\u00e4tsch G, Widmer C, Sch\u00f6lkopf B (2009) An empirical analysis of domain adaptation algorithms for genomic sequence analysis. Adv Neural Inform Process Syst 21:1433\u20131440","journal-title":"Adv Neural Inform Process Syst"},{"key":"11526_CR43","doi-asserted-by":"crossref","unstructured":"Guo H, Pasunuru R, Bansal M (2020) Multi-source domain adaptation for text classification via distancenet-bandits. In AAAI","DOI":"10.1609\/aaai.v34i05.6288"},{"key":"11526_CR44","unstructured":"Zhao S, Li B, Yue X, Gu Y, Xu P, Hu R, Chai H, Keutzer K (2019) Multi-source domain adaptation for semantic segmentation. NeurIPS"},{"issue":"8","key":"11526_CR45","doi-asserted-by":"publisher","first-page":"2274","DOI":"10.1109\/TMI.2023.3247543","volume":"42","author":"X Li","year":"2023","unstructured":"Li X, Lv S, Li M, Jiang Y, Qin Y, Luo H, Yin S (2023) SDMT: spatial dependence multi-task transformer network for 3d knee MRI segmentation and landmark localization. IEEE Trans Med Imaging 42(8):2274\u20132285. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1109\/TMI.2023.3247543","journal-title":"IEEE Trans Med Imaging"},{"issue":"3","key":"11526_CR46","doi-asserted-by":"publisher","first-page":"1958","DOI":"10.1109\/TII.2020.2993842","volume":"17","author":"X Li","year":"2020","unstructured":"Li X, Jiang Y, Li M, Yin S (2020) Lightweight attention convolutional neural network for retinal vessel image segmentation. IEEE Trans Ind Inf 17(3):1958\u20131967","journal-title":"IEEE Trans Ind Inf"},{"issue":"11","key":"11526_CR47","doi-asserted-by":"publisher","first-page":"3207","DOI":"10.1109\/TMI.2022.3181060","volume":"41","author":"K Hu","year":"2022","unstructured":"Hu K, Wu W, Li W, Simic M, Zomaya A, Wang Z (2022) Adversarial evolving neural network for longitudinal knee osteoarthritis prediction. IEEE Trans Med Imaging 41(11):3207\u20133217","journal-title":"IEEE Trans Med Imaging"},{"issue":"2","key":"11526_CR48","first-page":"1518","volume":"20","author":"Y Wan","year":"2023","unstructured":"Wan Y, Jiang Z (2023) Transcrispr: transformer based hybrid model for predicting CRISPR\/cas9 single guide RNA cleavage efficiency. IEEE Trans Med Imaging 20(2):1518\u20131528","journal-title":"IEEE Trans Med Imaging"},{"key":"11526_CR49","doi-asserted-by":"crossref","unstructured":"Manaka T, Van Zyl TL, Kar D (2022) Improving cause-of-death classification from verbal autopsy reports. arXiv:2210.17161","DOI":"10.1007\/978-3-031-22321-1_4"},{"key":"11526_CR50","doi-asserted-by":"crossref","unstructured":"Peters M, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L (2018) Deep contextualized word representations. NAACL","DOI":"10.18653\/v1\/N18-1202"},{"key":"11526_CR51","doi-asserted-by":"crossref","unstructured":"Boukkouri HE, Ferret O, Lavergne T, Noji H, Zweigenbaum P, Tsujii J (2020) Characterbert: Reconciling elmo and bert for word-level open-vocabulary representations from characters","DOI":"10.18653\/v1\/2020.coling-main.609"},{"key":"11526_CR52","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreita J, Jones L, Gomez AN (2017) Attention is all you need. NIPS, pp 6000\u20136010"},{"key":"11526_CR53","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","volume":"3","author":"K He","year":"2016","unstructured":"He K, Zhang X, Ren S, Jian S (2016) Deep residual learning for image recognition. AI Open 3:770\u2013778. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1109\/CVPR.2016.90","journal-title":"AI Open"},{"key":"11526_CR54","unstructured":"Ba LJ, Kiros JR, Hinton GE (2016) Layer normalization. arXiv:1607.06450"},{"key":"11526_CR55","unstructured":"Zhang X, Zhao J, LeCun Y (2015) Character-level convolutional networks for text classification. Adv Neural Inf Process Syst, pp 649\u2013657"},{"key":"11526_CR56","doi-asserted-by":"crossref","unstructured":"Verwimp L, Pelemans J, hamme HV, Wambacq P (2017) Character-word lstm language models. Proceedings of the 15th conference of the European chapter of the association for computational linguistics vol 1, pp 417\u2013427","DOI":"10.18653\/v1\/E17-1040"},{"key":"11526_CR57","unstructured":"Si Y, Roberts K (2018) A frame-based nlp system for cancer-related information extraction. AMIA Ann Symp Proc, pp 1524\u20131533"},{"key":"11526_CR58","doi-asserted-by":"crossref","unstructured":"Yan Z, Jeblee S, Hirst G (2019) Can character embeddings improve cause-of-death classification for verbal autopsy narratives? BioNLP@ACL","DOI":"10.18653\/v1\/W19-5025"},{"key":"11526_CR59","doi-asserted-by":"crossref","unstructured":"Affi M, Latiri C (2021) Be-blc: Bert-elmo-based deep neural network architecture for English named entity recognition task. Proc Comput Sci 192","DOI":"10.1016\/j.procs.2021.08.018"},{"key":"11526_CR60","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1016\/j.aiopen.2022.10.001","volume":"3","author":"T Lin","year":"2022","unstructured":"Lin T, Wang Y, Liu X, Qiu X (2022) A survey of transformers. AI Open 3:111\u2013132. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1016\/j.aiopen.2022.10.001","journal-title":"A survey of transformers. AI Open"},{"key":"11526_CR61","doi-asserted-by":"publisher","unstructured":"Guo M, Zhang Y, Liu T (2019) Gaussian transformer: a lightweight approach for natural language inference. In: Proceedings of AAAI, pp 6489\u20136496. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1609\/aaai.v33i01.33016489.","DOI":"10.1609\/aaai.v33i01.33016489."},{"key":"11526_CR62","doi-asserted-by":"publisher","unstructured":"Yang B, Tu Z, Wong DF, Meng F, Chao LS, Zhang T (2018) Modeling localness for self-attention networks. In: Proceedings of EMNLP. Brussels, Belgium, pp 4449\u20134458. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1109\/CVPR.2016.90.","DOI":"10.1109\/CVPR.2016.90."},{"key":"11526_CR63","doi-asserted-by":"crossref","unstructured":"Wang W, Li X, Ren H, Gao D, Fang A (2023) Chinese clinical named entity recognition from electronic medical records based on multisemantic features by using robustly optimized bidirectional encoder representation from transformers pretraining approach whole word masking and convolutional neural networks: model development and validation. JMIR Med Inform 11(e44597)","DOI":"10.2196\/44597"},{"key":"11526_CR64","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2021.103737","volume":"116","author":"J Kong","year":"2021","unstructured":"Kong J, Zhang L, Jiang M, Liu T (2021) Incorporating multi-level CNN and attention mechanism for Chinese clinical named entity recognition. J Biomed Inform 116:103737. https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1016\/j.jbi.2021.103737","journal-title":"J Biomed Inform"},{"key":"11526_CR65","unstructured":"Madabushi HT, Kochkina E, Castelle M (2020) Cost-sensitive BERT for generalisable sentence classification with imbalanced data. arXiv:2003.11563"},{"key":"11526_CR66","doi-asserted-by":"crossref","unstructured":"Wei JW, Zou K (2019) Eda: Easy data augmentation techniques for boosting performance on text classification tasks. arXiv:arXiv:1901.11196","DOI":"10.18653\/v1\/D19-1670"},{"key":"11526_CR67","unstructured":"Xiaoya L, Xiaofei S, Yuxian M, Junjun L, Fei W, Jiwei L (2020) Dice loss for data-imbalanced NLP tasks. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 465\u2013476"},{"key":"11526_CR68","first-page":"1","volume":"5","author":"TA Sorensen","year":"1948","unstructured":"Sorensen TA (1948) A method of establishing groups of equal amplitude in plant sociology based on similarity of species content and its application to analyses of the vegetation on Danish commons. Kong Dan Vidensk Selsk Biol Skr 5:1\u201334","journal-title":"Kong Dan Vidensk Selsk Biol Skr"},{"key":"11526_CR69","unstructured":"Abadi M, Agarwal A, Barham P, Brevdo E, Chen Z, Citro C, Corrado G.S, Davis A, Dean J, Devin M, Ghemawat S, Goodfellow I, Harp A, Irving G, Isard M, Jozefowicz R, Jia Y, Kaiser L, Kudlur M, Levenberg J, Man\u00e9 D, Schuster M, Monga R, Moore S, Murray D, Olah C, Shlens J, Steiner B, Sutskever I, Talwar K, Tucker P, Vanhoucke V, Vasudevan V, Vi\u00e9gas F, Vinyals O, Warden P, Wattenberg M, Wicke M, Yu Y, Zheng X (2015) TensorFlow: large-scale machine learning on heterogeneous systems. Software available from tensorflow.org. https:\/\/2.zoppoz.workers.dev:443\/https\/www.tensorflow.org\/"},{"key":"11526_CR70","doi-asserted-by":"publisher","first-page":"18","DOI":"10.12688\/gatesopenres.12812.1","volume":"2","author":"AD Flaxman","year":"2018","unstructured":"Flaxman AD, Harman L, Joseph J, Brown J, Murray CJ (2018) A de-identified database of 11,979 verbal autopsy open-ended responses. Gates Open Res 2:18","journal-title":"Gates Open Res"},{"key":"11526_CR71","unstructured":"Maas AL, Daly RE, Pham PT, Huang D, Ng AY, Potts C (2011) Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics"},{"key":"11526_CR72","unstructured":"Le QV, Mikolov T (2014) Distributed representations of sentences and documents. In Proceedings of the 31st international conference on machine learning (ICML 2014), pp 1188\u20131196"},{"key":"11526_CR73","unstructured":"Mtsamples (2022) Transcribed medical transcription sample reports and examples. Great collection of transcription samples. https:\/\/2.zoppoz.workers.dev:443\/https\/www.mtsamples.com\/"},{"key":"11526_CR74","doi-asserted-by":"crossref","unstructured":"Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L (2018) Deep contextualized word representations. arXiv:1802.05365","DOI":"10.18653\/v1\/N18-1202"},{"key":"11526_CR75","first-page":"37","volume":"37","author":"S Danso","year":"2013","unstructured":"Danso S, Johnson O, Ten Asbroek A, Soromekun S, Edmond K, Hurt C, Hurt L, Zandoh C, Tawiah C, Fenty J, Etego SA, Aygei SO, Kirkwood B (2013) A semantically annotated verbal autopsy corpus for automatic analysis of cause of death. ICAME J Int Comput Arch Modern Mediev English 37:37\u201369","journal-title":"ICAME J Int Comput Arch Modern Mediev English"},{"key":"11526_CR76","doi-asserted-by":"crossref","unstructured":"See A, Liu PJ, Manning CD (2017) Get to the point: Summarization with pointer-generator networks. arXiv:1704.04368","DOI":"10.18653\/v1\/P17-1099"},{"key":"11526_CR77","doi-asserted-by":"crossref","unstructured":"Jeblee S, Gomes M, Jha P, Rudzicz F, Hirst G (2019) Automatically determining cause of death from verbal autopsy narratives. BMC Med Inf Decis Mak 19(127)","DOI":"10.1186\/s12911-019-0841-9"},{"key":"11526_CR78","doi-asserted-by":"crossref","unstructured":"Jeblee S, Gomes M, Hirst G (2018) Multi-task learning for interpretable cause of death classification using key phrase predictions. In Proceedings of the BioNLP 2018 Workshop vol 34, no 19, pp 12\u201327","DOI":"10.18653\/v1\/W18-2302"},{"key":"11526_CR79","unstructured":"Manaka T, Van Zyl TL, Wade AN, Kar D (2022) Using machine learning to fuse verbal autopsy narratives and binary features in the analysis of deaths from hyperglycaemia. arXiv:2204.12169"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11526-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/article\/10.1007\/s11063-024-11526-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11526-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,15]],"date-time":"2024-07-15T07:23:38Z","timestamp":1721028218000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/link.springer.com\/10.1007\/s11063-024-11526-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,20]]},"references-count":79,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["11526"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1007\/s11063-024-11526-y","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2520073\/v1","asserted-by":"object"}]},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,20]]},"assertion":[{"value":"8 January 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"We wish to reaffirm that no known conflicts of interest related to this publication or substantial financial support might have impacted the research\u2019s findings.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"We further confirm that any aspect of the work covered in this manuscript that has involved either experimental animals or human patients has been conducted with the ethical approval of all relevant bodies and that such approvals are acknowledged within the manuscript (ethics clearance number: M110138).","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics Approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}}],"article-number":"177"}}